System for generating the results of organizing target data, information processing method, and program

JP7898131B1Active Publication Date: 2026-07-31ATLUS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ATLUS CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-31

AI Technical Summary

Benefits of technology

【0020】 本開示の発明によれば、対象データに含まれるテキスト情報を、元のテキストを直接復元可能なデータ構造を含まない特徴表現に変換し、当該特徴表現に基づいて編成結果を生成し、処理終了後に当該特徴表現を削除することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007898131000001_ABST
    Figure 0007898131000001_ABST
Patent Text Reader

Abstract

An information processing system that analyzes text information contained in multiple target data and generates an organized result in which the multiple target data are organized into predetermined units, Conversion unit, The Organization Department, Deletion control unit and Equipped with, The conversion unit converts the text information into a feature representation that does not include a data structure that allows the original text to be directly restored. The aforementioned organization unit generates the organization result based on the feature representation, The deletion control unit deletes the feature expression when triggered by a predetermined event related to the completion of processing. Information processing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention of the present disclosure relates to an information processing technology for generating a compilation result using text information included in a plurality of target data. More specifically, it relates to an information processing system, an information processing method, and a program that convert text information included in target data into a feature representation that does not include a data structure capable of directly restoring the original text, generate a compilation result based on the feature representation, and delete the feature representation after the processing ends. In one embodiment, the invention of the present disclosure is applied to lecture information submitted to an academic conference.

Background Art

[0002] In an academic conference, program compilation is performed to create a set of sessions that organize related lectures into certain groups based on the submitted lecture information. In recent years, there has been a demand for technologies that analyze lecture titles, lecture summaries, lecture transcripts, submission categories, keywords, etc. to group lectures and generate a compilation result.

[0003] On the other hand, since the lecture information submitted to an academic conference may include unpublished research content, consideration is also required in the handling of the data used in the processing. Conventionally, there are document analysis, clustering, and other related technologies, but there has not been enough technology that treats text information as a feature representation that does not include a data structure capable of directly restoring the original text, generates a compilation result based on the feature representation, and deletes the feature representation after the processing ends.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present invention aims to provide an information processing system, information processing method, and program that can convert text information contained in target data into a feature representation that does not include a data structure that allows the original text to be directly restored, generate an organization result based on the feature representation, and delete the feature representation after processing is complete. [Means for solving the problem]

[0006] An information processing system that analyzes text information contained in multiple target data and generates an organized result in which the multiple target data are organized into predetermined units, Conversion unit, The Organization Department, Deletion control unit and Equipped with, The conversion unit converts the text information into a feature representation that does not include a data structure that allows the original text to be directly restored. The aforementioned organization unit generates the organization result based on the feature representation, The deletion control unit deletes the feature expression when triggered by a predetermined event related to the completion of processing. Information processing system.

[0007] Further comprising an input control unit, The input control unit inputs only the text attributes necessary for generating the compilation result from the attributes included in the target data to the conversion unit, and excludes information that can directly identify an individual from the input to the conversion unit, in this information processing system.

[0008] Further equipped with a learning control unit, An information processing system in which the learning control unit controls the text information and feature representations not to be transmitted to an external learning platform, and to be used only for inference processing that generates the organization results, and not for additional training, fine tuning, or permanent knowledge storage of the model.

[0009] The conversion unit is configured to generate the feature representation that does not retain specific lexical information, word order information, and syntactic structure by converting the information extracted from the text information into numerical information including category probability values ​​or weighted scores corresponding to a plurality of predefined attribute items or abstract categories, thereby generating the feature representation that does not retain specific lexical information, word order information, and syntactic structure.

[0010] The conversion unit is configured to generate the feature representation by applying at least one process selected from principal component analysis, random projection, and low-dimensional mapping processing to the numerical information corresponding to the attribute item or the abstract category, and by performing compression that results in information loss.

[0011] The aforementioned conversion unit is configured not to permanently store the intermediate analysis information generated by the analysis using the language model, but to discard it after the analysis is complete. An information processing system comprising the intermediate analysis information including at least one of the token sequence, attention weight, hidden state, detailed log of assignment to the attribute item or the abstract category, and the result of individual application of the mapping dictionary.

[0012] The system further comprises an internal data store for storing the aforementioned feature representations and intermediate generated data generated during the organization process, The aforementioned internal data store is located in an internal network layer isolated from the external network, does not have an interface for receiving external queries, and is accessible only from authenticated internal services; it is an information processing system.

[0013] With an additional control server, The conversion unit and the organization unit are located within a controlled, isolated execution environment, are activated in response to requests from the control server via an internal API that is not publicly accessible, and do not accept direct activation by general users or direct connections to the internal data store; this is an information processing system.

[0014] It further comprises a user interface unit and a role-based access control unit, The user interface unit receives registration of the target data, receives execution instructions from the organization unit, and displays the organization results, but does not display the feature representation and the intermediate generated data. The role-based access control unit divides users into at least general users, administrators, and system administrators, and permits only the system administrators to view the feature representations and intermediate generated data, thereby providing an information processing system.

[0015] The aforementioned data is information processing data consisting of presentation information submitted to academic conferences, which may include unpublished research content, and which is complete data for each academic conference.

[0016] The aforementioned text attributes include at least one of the following: lecture title, lecture summary, lecture abstract, submission category, and keywords. The information processing system includes at least one of the following: author name, affiliation information, contact information, session chair information, and presenter identification information, which can directly identify the aforementioned individual.

[0017] The aforementioned organizational result is a collection of session units in the aforementioned academic conference, The deletion control unit deletes the feature representation, the structured metadata generated by the conversion unit, and the intermediate generated data generated by the organization unit, with respect to events related to the determination of the session unit set or the end of the academic conference as predetermined events. An information processing system that requires the re-execution of the analysis process based on the lecture information to retrieve the aforementioned organization result after deletion, and does not have a function to restore the organization result from the feature representation, the structured metadata, and the intermediate generated data.

[0018] An information processing method that analyzes text information contained in multiple target data and generates an organized result in which the multiple target data are organized into predetermined units, The conversion unit performs a conversion step of converting the text information into a feature representation that does not include a data structure that allows the original text to be directly restored, A compilation step in which a compilation unit generates the compilation result based on the feature expression, A deletion step in which a deletion control unit deletes the feature expression triggered by a predetermined event related to the end of processing An information processing method including the above.

[0019] A program for causing a computer to execute the information processing method.

Effect of the Invention

[0020] According to the invention of the present disclosure, text information included in target data can be converted into a feature expression that does not include a data structure capable of directly restoring the original text, a compilation result can be generated based on the feature expression, and the feature expression can be deleted after the processing is completed.

Brief Description of the Drawings

[0021] [Figure 1] A diagram showing the overall configuration of an information processing system according to an embodiment of the invention of the present disclosure [Figure 2] A flowchart showing the overall flow of information processing according to an embodiment of the invention of the present disclosure [Figure 3] A flowchart showing details of input control and feature expression generation processing according to an embodiment of the invention of the present disclosure [Figure 4] A flowchart showing details of compilation result generation processing according to an embodiment of the invention of the present disclosure [Figure 5] A flowchart showing details of deletion control and re-executability processing according to an embodiment of the invention of the present disclosure [Figure 6] A diagram showing an isolation execution and access control configuration according to an embodiment of the invention of the present disclosure

Mode for Carrying Out the Invention

[0023] In this specification, "Target Data" refers to each information unit that is processed in the information processing system or information processing method relating to the invention of this disclosure. Target Data may be, for example, each document in a document group, each proposal in a proposal information group, each research theme in a research theme information group, each examination subject in an examination subject information group, or other descriptive information units. In embodiments relating to academic conferences, the Target Data may be the lecture information corresponding to each lecture. In addition to the text information described later, the Target Data may also include identification information, management information, and other incidental information, but not all of them are necessarily subject to AI processing.

[0024] In this specification, "text information" refers to free-form text or similar descriptive information contained in the target data. "Text attribute" refers to descriptive information from the attributes contained in the target data that is input to the conversion unit for generating the compilation result. In this specification, when text attributes are treated as input for the conversion process, they may be recognized as text information. In embodiments relating to academic conferences, text information or text attributes may include, for example, at least a part of the lecture title, lecture summary, lecture abstract, submission category, keywords, etc. Here, the lecture summary and lecture abstract may be provided as separate input items, or they may be managed integrally as a single abstract field or a similar field. Furthermore, information provided by selection input or semi-structured input, such as submission category or keywords, may also be included in text information or text attributes in this specification, as long as it is descriptive information used for generating the compilation result.

[0025] In this specification, "information that can directly identify an individual" means identification attributes that are attached to or managed in conjunction with the target data, such as name, affiliation information, contact information, session chair information, presenter identification information, and other information that can directly identify a specific individual. In embodiments relating to academic conferences, information that can directly identify an individual may include, for example, author name, affiliation information, contact information, session chair information, presenter identification information, and other management information associated with an individual. On the other hand, descriptions of research content such as regional names, year, case attributes, disease names, etc., which may be included in the lecture summary or abstract, are not included in information that can directly identify an individual in this specification unless they are described in a manner that can directly identify a specific individual. The original lecture information DB41 may record such information that can directly identify an individual, but the input control unit 21 may exclude such information that can directly identify an individual from the information input to the conversion unit 22 for the generation of the compilation result.

[0026] In this specification, "structured metadata" refers to an intermediate semantic representation obtained by mapping text information to predefined attribute items or abstract categories. Structured metadata may be information organized along attribute axes such as subject area, research method, target entity, research objective, field classification, target audience, etc., and may include category labels, category IDs, category probability values, weighted scores, etc. Structured metadata is not the original text itself, nor does it differ from the feature representation described later. It is understood as an intermediate semantic representation that associates the semantic content of text information with predetermined attribute items or abstract categories.

[0027] In this specification, "numerical information" refers to numerical representations generated based on structured metadata. Numerical information may be composed of, for example, a vector, matrix, or similar numerical sequence that arranges categorical probability values, weighted scores, normalized values, or other numerical values ​​corresponding to multiple attribute items or abstract categories. Numerical information is a transformation of the semantic axes contained in the structured metadata into a form that is easy to process computationally, and forms the basis for generating feature representations described later.

[0028] In this specification, "feature representation" refers to a numerical representation obtained by applying compression processing to numerical information as necessary, which is used by the organization unit to generate the organization result. Compression processing can be, for example, principal component analysis, random projection, or low-dimensional mapping processing, but is not limited to these. In configurations where compression processing is omitted, the numerical information may be used directly as the feature representation. The feature representation may be a numerical representation that does not include a data structure that allows for the direct reconstruction of the original text. For example, it may be configured in a way that does not include the original text string itself, a token sequence, a sequence that preserves word order, a syntax tree, reference information to the original text fragment, or a reverse lookup table or other data structure for reconstruction of the original text. Therefore, the feature representation may be configured in a way that does not directly preserve specific lexical information, word order information, or syntactic structure, and is not the same concept as structured metadata. The former is an intermediate representation related to the organization of semantic content, while the latter is a numerical representation used for organization processing.

[0029] In this specification, “intermediate analysis information” means information temporarily generated during the process by a language model or equivalent analyzer to analyze text information. Intermediate analysis information may include, for example, token sequences, attention weights, hidden states, detailed logs of assignments to attribute items or abstract categories, and the results of individual application of mapping dictionaries. Intermediate analysis information is temporary information during the analysis process, distinct from structured metadata or feature representations themselves, and may, in one example, not be permanently stored after the completion of the analysis but may be discarded.

[0030] In this specification, "intermediate generated data" refers to various types of data generated during the organization process, in addition to structured metadata and feature representations. Intermediate generated data may include, for example, similarity matrices, cluster candidates, rearrangement candidates, evaluation values, and other data generated during the organization process. Unlike intermediate analysis information generated during the internal analysis process of a language model, intermediate generated data is understood as data generated during the execution of the organization process based on feature representations.

[0031] In this specification, "predetermined unit" refers to a unit for grouping target data. The predetermined unit can be set as appropriate depending on the application, and may be, for example, a session, a group of themes, a group of classifications, a group of reviews, or other collective units. In embodiments relating to academic conferences, the predetermined unit may be a session.

[0032] In this specification, "organization result" refers to result information that includes the correspondence between target data and predetermined units. The organization result may include assignment results indicating which predetermined unit each target data belongs to, and may also include, as necessary, the name of the predetermined unit, display order, evaluation information, and other supplementary information. In embodiments relating to academic conferences, the organization result can be understood as a set of session units, each of which presentation information is associated with a specific session.

[0033] In this specification, "external learning infrastructure" refers to a group of external servers or services that perform additional training, fine-tuning, or permanent knowledge storage of a model. In contrast, the inference execution infrastructure, which is located within the isolated execution environment 30 described later and performs inference processing for generating the organization results, is distinguished from the external learning infrastructure. Therefore, even if the same or similar model format is used, a configuration used exclusively for inference within the isolated execution environment 30, without the purpose of additional training, etc., is not included in the definition of external learning infrastructure in this specification.

[0034] In this specification, “a predetermined event related to the termination of processing” means an event that initiates the deletion of feature representations, structured metadata, or intermediate generated data. In embodiments relating to academic conferences, the predetermined event may be, for example, the confirmation of a set of sessions or the termination of the academic conference, but is not limited to these, and may also include the expiration of a predetermined retention period, an administrator's instruction, or the fulfillment of other termination conditions.

[0035] The specific meanings of the above terms will be appropriately defined according to the embodiments described later, and are not limited to specific examples of the inventions disclosed herein. Furthermore, even when terms related to academic conferences are used in this specification, this is merely an explanation of one embodiment for the sake of clarity, and the inventions disclosed herein can be understood as a general information processing technology that securely handles text information contained in multiple data sets and generates a compilation result in which the multiple data sets are organized into predetermined units.

[0036] The embodiments of the invention disclosed herein will be described below with reference to Figures 1 to 6. The invention disclosed herein is widely applicable as an information processing technology that securely handles text information contained in multiple target data and generates an organized result in which the multiple target data are organized into predetermined units. As an example, the case in which presentation information submitted to an academic conference is used as target data will be described below.

[0037] Presentation information submitted to academic conferences may include text attributes such as presentation title, abstract, summary, submission category, and keywords, but it may also include unpublished research content, and it is transient data that is completed for each conference. Therefore, when analyzing this presentation information to generate session-based sets, it is not sufficient to simply perform a similarity-based organization process. It is desirable to have a structure that also takes into account the scope of input information, the nature of the data used for analysis, the method of retaining the data generated during the analysis process, and the method of data management after processing is complete.

[0038] Therefore, the information processing system 1 according to this embodiment is configured as an information processing platform that, in addition to a function for generating a compilation result, includes input control that takes in only the text attributes necessary for generating the compilation result, dedicated use of such text attributes for inference without using them for additional learning, deletion control that deletes feature representations, etc. after processing is complete, isolated execution that places the analysis function in an environment partitioned from the outside, and access control that restricts who can access the intermediate data. In other words, the core of the invention of this disclosure is not limited to the lecture information compilation algorithm itself, but lies in the platform configuration for generating a compilation result while safely handling the target data.

[0039] Figure 1 is a schematic diagram showing the relationships between the components of the information processing system 1 according to this embodiment. As shown in Figure 1, the information processing system 1 includes a user terminal 10, a control server 20, an isolated execution environment 30, an internal API 31, an internal data store 40, a source lecture information DB 41, and a compilation result storage area 42. Each of these components may be implemented by consolidating them into a single physical device, or they may be implemented by distributing them across multiple physical devices or virtual computing resources.

[0040] The user terminal 10 is a terminal used for receiving registration of target data, receiving execution instructions, and displaying the organization results. It accepts user input and can display various information provided by the control server 20. The user terminal 10 and the control server 20 are connected in a communication manner, and together they realize the functions of a user interface unit for registering target data, starting the organization process, and viewing the generated organization results. In one example, the user interface unit may be configured to receive registration of target data, receive execution instructions from the organization unit 23, and display the organization results, but may not have functions for displaying or acquiring structured metadata, feature representations, and intermediate generated data.

[0041] The control server 20 functions as the control entity in this embodiment and functionally includes at least an input control unit 21, a learning control unit 24, a deletion control unit 25, and a role-based access control unit 50. The input control unit 21 is responsible for selecting text attributes necessary for generating the compilation result from the lecture information obtained from the original lecture information DB 41. The learning control unit 24 is responsible for controlling that the data used in the conversion process and compilation process described later be used only for inference processing. The deletion control unit 25 is responsible for controlling the deletion of feature representations, structured metadata, and intermediate generated data in response to predetermined events related to the completion of processing. The role-based access control unit 50 is responsible for controlling access rights to various data handled within the system.

[0042] Within the isolated execution environment 30, a conversion unit 22 and an organization unit 23 are arranged. The conversion unit 22 is responsible for analyzing the text information selected by the input control unit 21 and generating structured metadata and feature representations. The organization unit 23 is responsible for generating organization results in which the target data is organized into predetermined units based on the feature representations obtained by the conversion unit 22. In one example, the control server 20 may start or control the conversion unit 22 and the organization unit 23 within the isolated execution environment 30 via an internal API 31. In this way, by arranging the main functions related to analysis and organization within the isolated execution environment 30, these processes can be executed in an environment partitioned from the outside.

[0043] The original presentation information DB41 may store presentation information corresponding to each presentation submitted to the academic conference. Presentation information may include, for example, text attributes such as presentation title, presentation abstract, presentation summary, submission category, and keywords, as well as author name, affiliation information, contact information, and other information or management information that can directly identify an individual. The presentation abstract and presentation summary may be stored as separate items, or they may be stored as a single abstract field or a similar field. However, not all information recorded in the original presentation information DB41 is used for AI processing. The input control unit 21 selects only the attributes necessary for generating the compilation result and excludes unnecessary information, especially information that can directly identify an individual, from the AI ​​processing target. On the other hand, descriptions of research content included in the presentation abstract or presentation summary may be used as text attributes necessary for generating the compilation result, as long as they do not themselves constitute information that can directly identify an individual.

[0044] The internal data store 40 is a storage area for internal use of the transformation and organization processes, and can store structured metadata and feature representations generated by the transformation unit 22, as well as various intermediate generated data generated by the organization unit 23. Intermediate generated data may include, for example, similarity matrices, cluster candidates, rearrangement candidates, evaluation values, and other data used in the organization process. The transformation unit 22 may write the generated structured metadata and feature representations to the internal data store 40 and refer to them as needed. The organization unit 23 may also perform the organization process by referring to the feature representations stored in the internal data store 40, and may write or refer to the intermediate generated data generated in the process to the internal data store 40. In other words, the internal data store 40 functions as a storage base that can be used internally by the transformation unit 22 and the organization unit 23.

[0045] The final organization result generated by the organization unit 23 is stored in the organization result storage area 42. The organization result stored in the organization result storage area 42 may be displayed on the user terminal 10 via the control server 20, and may be applied to other operational processing or display functions as needed. Thus, in this embodiment, the original lecture information DB 41 functions as the source of the target data, the conversion unit 22 and organization unit 23 in the isolated execution environment 30 are responsible for conversion and organization based on the target data, the internal data store 40 functions as the storage location for internal processing data, and the organization result storage area 42 functions as the storage location for the final output.

[0046] As described above, in the information processing system 1 according to this embodiment, a series of relationships—acquisition of target data, selection of necessary attributes, conversion to structured metadata and feature representations, generation of the organization result, storage of internal data and the final result, and display of the organization result—are realized through the cooperation of the control server 20, isolated execution environment 30, internal data store 40, original presentation information DB 41, and organization result storage area 42. The specific operation details of each of these configurations will be described later with reference to Figures 2 to 6.

[0047] Figure 2 is a flowchart showing the main flow of the information processing system 1 according to this embodiment. As shown in Figure 2, the information processing according to this embodiment may include: step S1 of acquiring target data; step S2 of selecting necessary text attributes and excluding information that can directly identify an individual; step S3 of analyzing the text information to generate structured metadata and feature representations; step S4 of generating an organization result based on the feature representations; step S5 of saving or applying the organization result; step S6 of deleting feature representations, structured metadata, and intermediate generated data in response to the occurrence of a predetermined event; and step S7 of re-executing at least steps S2 to S5 based on the original presentation information when the organization result is to be reacquired. Each step will be described in order below.

[0048] In step S1, the target data is acquired. In this embodiment, presentation information submitted to an academic conference is used as the target data, and in one example, presentation information corresponding to each presentation may be acquired from the original presentation information DB41. However, the invention of this disclosure is not limited thereto, and the target data may be received via a network, or read from an external recording medium, another management system, or other storage area.

[0049] In step S2, the text attributes necessary for generating the compilation results are selected from the acquired target data, and information that can directly identify individuals is excluded. In other words, in this embodiment, even if the target data contains various types of information, not all of it is uniformly input to the AI ​​processing, the input control unit 21 extracts only the attributes necessary for generating the compilation results and excludes unnecessary attributes, especially information that can directly identify individuals, from the AI ​​processing target. However, descriptions of research content included in the lecture summary or abstract can be used as text attributes necessary for generating the compilation results, as long as they do not themselves constitute information that can directly identify individuals. Details of step S2 will be explained with reference to Figure 3, which will be described later.

[0050] In step S3, the text information selected in step S2 is analyzed to generate structured metadata and feature representations. In this embodiment, the text information is analyzed by a language model or a similar analysis model, and structured metadata is obtained by associating it with predefined attribute items or abstract categories. Furthermore, feature representations used in the organization process are generated from numerical information based on the structured metadata. Details of step S3 will be explained later with reference to Figure 3.

[0051] In step S4, the organization results are generated based on the feature representations. For example, the organization unit 23 may group the target data sets based on the similarity of the feature representations for each target data set and generate a predetermined unit, such as a set for each session in an academic conference. Details of the specific processing related to step S4 will be explained later with reference to Figure 4.

[0052] In step S5, the generated organization result is saved or applied. Here, "save or apply" is a broad concept that is not limited to saving the organization result in the organization result storage area 42, but may also include, for example, reflecting it in operational data, reflecting it in display data, or passing it on to other program organization processes. Therefore, the organization result may not only be saved, but may also be used for subsequent processes necessary for operation.

[0053] In step S6, feature representations, structured metadata, and intermediate generated data are deleted in response to the occurrence of a predetermined event related to the completion of processing. The predetermined event in this embodiment may be, for example, the confirmation of a session-based set or the end of an academic conference. This prevents the unnecessary retention of internal data generated during or for the purpose of the organization process. The specific processing details of step S6 will be explained with reference to Figure 5, which will be described later.

[0054] In step S7, if it is necessary to re-obtain the organization results, at least steps S2 to S5 are re-executed based on the original presentation information. What is important here is that the re-obtaining is not a process of restoring the organization results from deleted feature representations, structured metadata, or intermediate generated data. In other words, in this embodiment, instead of reproducing the organization results from deleted internal data, the system is configured to re-execute the selection of necessary text attributes, generation of feature representations, generation and saving or application of the organization results based on the original presentation information recorded in the original presentation information DB41.

[0055] The main flow shown in Figure 2 above represents the entry point to the overall processing in the information processing system 1 according to the present invention, and each detailed process shown in Figures 3 to 6 is a concrete implementation of the corresponding part of steps S2 to S7. Below, the input control and feature representation generation processes will be explained with reference to Figure 3.

[0056] Figure 3 is a flowchart detailing the input control and feature representation generation process according to this embodiment. As shown in Figure 3, this process may include the steps of: S2-1 acquiring presentation information from the original presentation information DB41; S2-2 extracting text attributes necessary for generating the compilation result; S2-3 excluding information that can directly identify individuals; S3-1 analyzing the text information using a language model; S3-2 mapping the information to predefined attribute items or abstract categories to generate structured metadata; S3-3 generating numerical information including category probability values ​​or weighted scores; S3-4 performing compression processing as necessary; S3-5 generating a feature representation that does not include a data structure that can directly restore the original text; and S3-6 discarding intermediate analysis information.

[0057] In step S2-1, presentation information is retrieved from the original presentation information DB41. The presentation information may include text attributes that can be used for organization, such as presentation title, presentation abstract, presentation summary, submission category, and keywords, as well as author names, affiliation information, contact information, session chair information, presenter identification information, and other information that can directly identify an individual or management information. In other words, the original presentation information DB41 may hold not only the information necessary for the organization process, but also various information necessary for operational management.

[0058] In step S2-2, text attributes necessary for generating the compilation result are extracted from the lecture information. In this embodiment, the lecture title, lecture summary, lecture abstract, submission category, and keywords may be used as the extraction targets. Note that the lecture summary and lecture abstract may be managed as separate input items, or they may be managed as a single abstract field or a similar field. Furthermore, even if the submission category or keywords are provided by selection input or semi-structured input, they may be treated as text attributes in this embodiment as long as they are attributes used for generating the compilation result.

[0059] In step S2-3, information that can directly identify an individual, such as author name, affiliation information, contact information, session chair information, and presenter identification information, is excluded. Importantly, while the original lecture information DB41 itself may contain information that can directly identify an individual, the input control unit 21 does not include this information in the text information input to the language model or other analyzers. In other words, in this embodiment, even if the original data contains information that can directly identify an individual, the AI ​​input for generating the compilation result does not include such information. On the other hand, descriptions of research content included in lecture abstracts, etc., can be treated as text attributes for generating the compilation result, as long as the description itself does not constitute information that can directly identify an individual. This structurally suppresses the inclusion of information that can directly identify an individual, which is unnecessary for the compilation process, into the AI ​​processing target.

[0060] In step S3-1, text information is analyzed using a language model. The language model may be a large-scale language model, an encoder-type language model, a classification model, or any other model with text analysis capabilities; its specific implementation is not limited. For example, this analysis process may be performed exclusively for inference by the conversion unit 22 within the isolated execution environment 30, under the control of the learning control unit 24. That is, this process is performed not for additional learning, fine-tuning, or permanent knowledge accumulation, but as an inference process necessary for generating the current organization result.

[0061] In step S3-2, based on the analysis results, text information is mapped to predefined attribute items or abstract categories to generate structured metadata. Examples of attribute items or abstract categories include subject area, research method, target entity, research objective, field classification, and target audience. In embodiments relating to academic conferences, these semantic perspectives may be defined as multiple semantic axes; for example, each presentation information may be organized along approximately 12 semantic axes. However, the number, names, and granularity of attribute items or abstract categories are not fixed and can be changed as appropriate depending on the target field, operational policy, or organizational purpose.

[0062] Structured metadata is an intermediate semantic representation that may include category labels, category IDs, category probability values, weighted scores, etc. For example, for a given lecture, a degree of relevance or contribution may be assigned to each of multiple attribute items or abstract categories, and these may be organized in a predetermined data format. Here, structured metadata is neither the original lecture title or abstract itself nor the feature representation described later. In other words, structured metadata is positioned as an intermediate representation that organizes text information along semantic axes.

[0063] In step S3-3, numerical information including category probability values ​​or weighted scores is generated based on structured metadata. For example, numerical values ​​corresponding to multiple attribute items or abstract categories may be arranged into a vector or matrix. In one example, a weight according to the operational policy may be set for each attribute item or abstract category, and this weight may be reflected in the category probability value or weighted score. This allows the semantic characteristics of each presentation information to be converted into a format suitable for computational processing. Note that the method of generating numerical information is not limited to this and may include normalization, weight correction, score integration, and other computational processing.

[0064] In steps S3-4, if necessary, at least one compression process selected from principal component analysis, random projection, and low-dimensional mapping is applied to the numerical information. In one embodiment, this compression process functions as a compression with information loss, reducing the dimensionality of the numerical information, suppressing redundancy, and converting it into a representation suitable for the organization process. As a result, the feature representation becomes a numerical representation that is further abstracted from the representation of the original text information. Note that if the dimensionality or representation format of the numerical information is already suitable for the organization process, the compression process may be omitted.

[0065] In step S3-5, the conversion unit 22 expresses the information extracted from the text information as numerical information including category probability values ​​or weighted scores corresponding to a plurality of predefined attribute items or abstract categories, and applies compression processing as necessary to generate a feature representation used by the organization unit 23 to generate the organization result. The feature representation in this embodiment is a numerical representation obtained from numerical information or numerical information after compression processing, and does not include a data structure that can directly restore the original text. Specifically, the feature representation does not include the original text string, token sequence, sequence information that preserves word order, syntax tree, or reverse lookup information for original text restoration, and is configured as a vector representation for comparing semantic proximity or relationship between target data.

[0066] Here, structured metadata and feature representations are clearly distinguished. That is, structured metadata is an intermediate semantic representation obtained by mapping text information to attribute items or abstract categories, while feature representations are numerical representations generated based on the structured metadata and, after undergoing compression processing as necessary, are used by the organization unit 23 to generate the organization result. Therefore, in this embodiment, the stage of organizing the semantic content and the stage of generating representations used in the organization calculation are distinguished.

[0067] In step S3-6, intermediate analysis information is discarded. Intermediate analysis information may include, for example, token sequences, attention weights, hidden states, detailed logs of assignments to attribute items or abstract categories, and individual application results of mapping dictionaries. This intermediate analysis information is temporarily generated during the analysis process by the language model and is not subject to permanent storage; it is discarded after the generation of structured metadata or feature representations is complete. In this embodiment, by not permanently storing the intermediate analysis information in the internal data store 40, the continuous retention of the detailed analysis process itself can be suppressed.

[0068] According to the process shown in Figure 3 above, even if the original lecture information DB 41 contains raw data that can directly identify individuals, the input control unit 21 selects only the text attributes necessary for generating the compilation result, and the information that can directly identify individuals is excluded from the AI ​​processing target. On the other hand, descriptions of research content included in the lecture summary or abstract can be processed as part of the text attributes, as long as they do not themselves constitute information that can directly identify individuals. Furthermore, the text information is progressively converted into structured metadata and feature representations, and intermediate analysis information is not permanently stored. For this reason, this embodiment can function as an information processing platform that performs the analysis processing necessary for generating the compilation result while simultaneously achieving secure input control, organization of semantic representations, generation of numerical representations for compilation, and non-retention of analysis process information.

[0069] For example, if the presentation information submitted to a certain academic conference includes the presentation title "Damage Determination Method Using Bridge Inspection Images," the presentation abstract "Evaluating Bridge Damage Progression Based on Image Features and Time-Series Changes," the submission category "Structural Engineering," the keywords "Bridge, Damage Determination, Image Analysis," and the author's name and affiliation information, the input control unit 21 may exclude the author's name and affiliation information and input the presentation title, presentation abstract, submission category, and keywords to the conversion unit 22. Based on the input, the conversion unit 22 may generate structured metadata corresponding to attribute items or abstract categories such as the subject area "Structural Engineering," the research method "Image Analysis," the target entity "Bridge," the research objective "Damage Determination," and the target audience "Maintenance and Management Field," generate numerical information from the structured metadata, and generate a feature representation by applying compression processing as necessary.

[0070] Figure 4 is a flowchart detailing the organization result generation process according to this embodiment. As shown in Figure 4, this process may include the steps of: acquiring feature representations (S4-1); grouping presentation groups based on the similarity of the feature representations (S4-2); reintegrating outliers, re-evaluating, readjusting, or adjusting the number of items (S4-3) as needed; generating a set of session units (S4-4); generating proposed session titles as needed (S4-5); and saving the organization results and applying them to operational data as needed (S4-6).

[0071] In step S4-1, the organization unit 23 obtains a feature representation to be used in the organization process. The feature representation may be, for example, a numerical representation corresponding to each presentation information stored in the internal data store 40, and the organization unit 23 may obtain the feature representation by referring to the internal data store 40. If necessary, the organization unit 23 may also refer to structured metadata or organization condition information in addition to the feature representation.

[0072] In step S4-2, the organization unit 23 groups the presentations based on the similarity of their feature expressions. That is, the organization unit 23 evaluates the proximity or relationship between the feature expressions corresponding to each presentation and groups semantically similar presentations into the same candidate group, thereby forming the groups that form the basis for generating the organization result. What is important in this embodiment is that the organization unit 23 generates the organization result based on feature expressions, rather than the original presentation titles or abstracts themselves.

[0073] In one example, the organization unit 23 may perform the grouping process in AI-driven mode. In AI-driven mode, the similarity of feature representations is evaluated across all presentations submitted to the academic conference, and presentation groups may be formed based on the semantic proximity of the presentations to each other, without being overly constrained by existing submission categories. According to such a mode, session candidates that reflect cross-curricular subject groups or potential relationships that were not explicitly stated at the time of submission can be generated.

[0074] In another example, the organization unit 23 may perform grouping processing in a submission category-focused mode. In submission category-focused mode, presentation groups may be generated using existing submission categories as the basic unit. For example, cluster generation may be performed within each submission category based on the similarity of feature representations, or complementary reorganization may be performed between adjacent or related submission categories. Such a mode makes it easier to maintain consistency with the classification system at the time of submission and to ensure ease of understanding in operation or compatibility with existing operations. Note that the AI-driven mode and the submission category-focused mode are both specific examples of the organization unit 23, and the basic configuration of the invention disclosed herein is not limited to these modes.

[0075] In step S4-3, outlier reintegration, reevaluation, readjustment, or adjustment of the number of presentations may be performed as needed. For example, the initially formed groups of presentations may be reevaluated from the perspectives of internal consistency, intergroup segregation, operational validity, and other factors, and presentations with relatively low suitability may be temporarily moved to a separate group, after which their possibility of being reassigned to other groups of presentations may be reevaluated. Alternatively, outlier reintegration may be performed to reintegrate small groups or isolated presentations resulting from the initial grouping into other similar groups of presentations.

[0076] Furthermore, as a readjustment or adjustment of the number of presentations, the presentation groups may be reorganized or the presentations rearranged, taking into account the number of presentations per session for each desired presentation format, upper or lower limits on the number of sessions, operationally required balance conditions, and other constraints. For example, if the number of presentations in a certain group exceeds a desired range, the group may be split, and if it falls below that range, it may be merged with or rearranged with other adjacent presentation groups. As specific implementations of grouping, re-evaluation, or readjustment, density-based methods, affinity-based methods, constrained optimization methods, and other known methods can be used as appropriate, but the core of the invention of this disclosure is not in these specific algorithms themselves, but in realizing the generation of organization results based on feature representation on a secure processing platform.

[0077] In step S4-4, a set of session units is generated based on the results of grouping and necessary readjustments. That is, the organizing unit 23 associates each presentation information with one of the sessions and, if necessary, may generate an organizing result that includes a list of presentations belonging to each session, session identifiers, session order, and other supplementary information. In this embodiment, the set of session units obtained in this way corresponds to the "organizing result".

[0078] In steps S4-5, if necessary, proposed session titles are generated for a set of session units. These proposed session titles may be generated, for example, based on the presentation titles, keywords, structured metadata, or representative features of the presentations belonging to the session. This makes it easier for the organizers to understand the content of the formed group of sessions. However, the generation of proposed session titles is merely one example of an auxiliary function of the organization result and does not limit the essence of the invention of this disclosure.

[0079] In step S4-6, the generated organization results are saved and applied to the operational data as needed. For example, the organization unit 23 or the control server 20 may save the organization results to the organization result storage area 42, or it may apply them as display data, operational management data, or data passed to subsequent program organization processing. In addition, similarity matrices, cluster candidates, rearrangement candidates, evaluation values, and other intermediate generated data generated during the organization process may be recorded in the internal data store 40 as needed.

[0080] According to the process shown in Figure 4 above, the organization unit 23 can group presentations based on feature representations and generate a set of session units after necessary readjustments. Furthermore, the AI-driven mode, submission category-focused mode, outlier reintegration, number condition adjustment, and session title proposal generation can all be adopted as specific embodiments of the organization unit 23. On the other hand, the core of this embodiment lies not in these specific algorithms themselves, but in generating the organization results on a secure information processing infrastructure that includes input control, feature representation, internal storage, and deletion control, which will be described later.

[0081] Figure 5 is a flowchart detailing the deletion control and re-execution process according to this embodiment. As shown in Figure 5, the process may include: step S6-1 detecting an event related to the determination of a session-based set or the end of an academic conference; step S6-2 identifying the data to be deleted; step S6-3 deleting the feature representation, structured metadata, and intermediate generated data; step D6-1 determining whether it is necessary to re-acquire the organization results after the deletion; step S6-4 not restoring from the deleted data; and step S6-5, if the organization results are to be re-acquired, re-executing at least the text attribute selection, feature representation generation, organization result generation, and saving or applying processes based on the original presentation information recorded in the original presentation information DB41.

[0082] In step S6-1, the deletion control unit 25 detects a predetermined event related to the completion of processing. In this embodiment, the predetermined event may be, for example, the confirmation of a session-based collection or the end of an academic conference. In addition, in one example, the expiration of a predetermined retention period, a deletion instruction by an administrator, the fulfillment of operationally defined termination conditions, or other events may be used as predetermined events.

[0083] In step S6-2, the data to be deleted is identified in response to the occurrence of a predetermined event. In this embodiment, the data to be deleted may include feature representations and structured metadata, as well as intermediate generated data generated during the organization process. Intermediate generated data may include, for example, similarity matrices, cluster candidates, rearrangement candidates, evaluation values, and other data used during the organization process. In other words, in this embodiment, internal data used in the preliminary or auxiliary determination of the organization result itself can be managed as data to be deleted.

[0084] In step S6-3, the deletion control unit 25 deletes the identified items to be deleted. For example, the deletion control unit 25 may delete feature representations, structured metadata, and intermediate generated data recorded in the internal data store 40, and may also treat corresponding data remaining in working memory, cache area, temporary file area, or other temporary storage area as items to be deleted, if necessary. This prevents unnecessary retention of internal data generated for the team organization process after the tournament has ended or the team organization has been finalized.

[0085] In step S6-4, no restoration is performed from deleted data. In other words, this embodiment does not have a function to restore the organization result from deleted feature representations, structured metadata, and intermediate generated data. To put it another way, deletion is not merely a display-level invisibility; the system is configured not to provide a function to regenerate the organization result based on the deleted internal data after deletion. This makes it possible to operate without relying on deleted internal data.

[0086] In step S6-5, if it becomes necessary to re-acquire the organization results after deletion, the following processes are performed again, at least text attribute selection, feature representation generation, organization result generation, and saving or application, based on the original presentation information recorded in the original presentation information DB41. In other words, re-acquisition is performed by re-analysis from the original data, not by restoration from deleted data. For example, the input control unit 21 may re-select the necessary text attributes, the conversion unit 22 may re-generate structured metadata and feature representations, and the organization unit 23 may re-generate session-based sets.

[0087] In this embodiment, retrieving the organization results after deletion requires re-execution of each process, and does not immediately obtain the organization results by re-reading the deleted feature representations, structured metadata, or intermediate generated data. Therefore, when retrieving, the organization results may be generated again, reflecting the organization conditions, operating conditions, or processing settings that are in effect at that time. Such a configuration is suitable for applications that handle transient data that is completed for each academic conference, and it is easy to suppress information mixing between conferences or unintended carrying over of internal data.

[0088] According to the process shown in Figure 5 above, in this embodiment, feature representations, structured metadata, and intermediate generated data are deleted in accordance with predetermined events related to the completion of processing, and the compilation results can be retrieved by re-analyzing the original presentation information recorded in the original presentation information DB41 when necessary. In other words, this embodiment does not permanently retain the internal data necessary for generating the compilation results, but can function as an information processing platform that ensures the possibility of retrieval after deletion by re-executing from the original data.

[0089] Furthermore, in this embodiment, security can be further enhanced by appropriately controlling the arrangement of the conversion unit 22 and the organization unit 23, the access method to the internal data store 40, and the relationship with the external learning platform, in conjunction with deletion control. These isolation execution and access control configurations will be described next with reference to Figure 6.

[0090] Figure 6 shows the isolation execution and access control configuration according to this embodiment. Figure 6 mainly simplifies the relationships between the components, and the details of the actual communication control, execution policy, authentication conditions, and access permission conditions are not limited to those shown in the figure. Therefore, in this embodiment, in addition to the arrangement of each component shown in Figure 6, matters such as not transmitting to the external learning platform, using it exclusively for inference, the non-exposure of the internal API 31 to the outside, the non-acceptance of external queries by the internal data store 40, and the restriction of the viewing entity by the role-based access control unit 50 are supplemented by text.

[0091] First, in this embodiment, "external learning infrastructure" refers to a group of external servers or services that perform additional training, fine-tuning, or permanent knowledge storage of the model. In contrast, the transformation unit 22 located within the isolated execution environment 30, or the execution infrastructure for the language model or analysis model used by the transformation unit 22, is positioned as an inference execution infrastructure that performs inference processing for generating the organization result, and is distinguished from the external learning infrastructure. That is, even if the same or similar model format is used, the infrastructure that performs additional training, etc., and the infrastructure that processes data specifically for inference on the input data at that time are distinguished functionally and operationally.

[0092] The learning control unit 24 is responsible for control functions to make the distinction effective. Specifically, the learning control unit 24 controls the communication path, connection destination, available external services, or execution policy so that text information, structured metadata, and feature representations are not transmitted to an external learning platform, and so that this data is used only for inference processing for generating the organization result. In other words, in this embodiment, the text information extracted from the lecture information to be organized, and the structured metadata and feature representations generated based thereon, are not used for additional learning, fine tuning, or permanent knowledge storage, but are processed for the limited purpose of generating the organization result related to the conference.

[0093] The control by the learning control unit 24 may be implemented, for example, by a configuration that limits the communication destinations reachable from the isolated execution environment 30 to the internal system, by a configuration that applies an execution policy that does not permit calls to the learning API or external storage function, or by a configuration that rejects processing commands that involve external transmission. Alternatively, the processing library, model execution environment, or service call settings used by the conversion unit 22 or the organization unit 23 may be configured to disable the learning mode and only permit the inference-only mode. This prevents the data processed internally for generating the organization result from being stored or reused beyond its intended purpose.

[0094] The internal data store 40 is a storage infrastructure located in the internal network layer, isolated from the external network. The internal data store 40 does not have an external query acceptance interface and is configured to be accessible only from authenticated internal services. For example, the internal services may include the transformation unit 22 and the organization unit 23 within the isolated execution environment 30, or internal processing functions associated with them. In other words, the internal data store 40 does not accept queries directly from user terminals 10, unspecified external systems, or normal management screens.

[0095] Since the internal data store 40 can record structured metadata, feature representations, and intermediate generated data, it is desirable that access to it be limited. For this reason, in this embodiment, there is no direct connection path from the user terminal 10 to the internal data store 40, and even access from the management terminal 11 is routed through the control server 20 and the role-based access control unit 50. The conversion unit 22 and the organization unit 23 are activated from the control server 20 via the internal API 31 and refer to or update the internal data store 40 as needed. This limits the access path to the internal data store 40 to the internal processing system.

[0096] The conversion unit 22 and the organization unit 23 are located within an isolated execution environment 30. The isolated execution environment 30 may be implemented as, for example, an on-premises environment, a virtual private cloud environment, a closed configuration within a single cloud account, a container execution environment, or other managed execution environment, but in any case it is configured as a partitioned execution area that is not directly accessible by external users. The control server 20 starts the conversion unit 22 and the organization unit 23 or transmits processing requests via an internal API 31 that is not publicly disclosed to the outside world. Here, the internal API 31 is not a public API for general users or external businesses, but a private interface used for control communication between internal components.

[0097] Therefore, general users and administrators cannot directly activate the conversion unit 22 and the organization unit 23, nor can they directly send arbitrary processing requests or arbitrary data reference requests to them. Operations accepted from the user terminal 10 are limited to registration acceptance, execution instruction acceptance, and organization result viewing via the control server 20, and do not permit direct manipulation of the internal state, internal parameters, structured metadata, feature representations, or intermediate generated data of the conversion unit 22 and the organization unit 23. This allows the activation and reference paths for internal processing to be consolidated with the control server 20 and the internal API 31, clearly defining the restrictions on who can access the system.

[0098] The user interface functions implemented by the user terminal 10 and the control server 20 are limited to receiving registration of target data, receiving execution instructions, and displaying the organization results. In other words, the normal user interface in this embodiment does not have the function of displaying or acquiring structured metadata, feature representations, and intermediate generated data. For example, the user terminal 10 may display the final organization results, such as which session each presentation is associated with and which presentations belong to each session, but may not display the similarity matrix, cluster candidates, rearrangement candidates, evaluation values, and other internal data that are internally generated during the process of generating the organization results.

[0099] Limiting user interface functionality is not merely about simplifying display items; it also has structural significance in restricting who can access intermediate data. Specifically, by not including functions for viewing, extracting, downloading, or reusing structured metadata, feature representations, and intermediate generated data in the screens or operating systems used daily by general users and administrators, it is possible to suppress the circulation of such data within the scope of normal operations. Such a configuration is effective in clearly defining the data as internal data for generating the compilation results and limiting its handling to the bare minimum necessary.

[0100] The role-based access control unit 50 controls the entities that access each component and data according to their roles. In this embodiment, users are classified into at least general users, administrators, and system administrators. General users are, for example, those who register lecture information or view the results of the organization process, while administrators are those who give instructions for the organization process, confirm the results of the organization process, or make operational adjustments. In contrast, system administrators are entities with administrative authority who can refer to the internal state to the extent necessary for system maintenance, auditing, or troubleshooting.

[0101] The role-based access control unit 50, based on its classification, grants access to structured metadata, feature representations, similarity matrices, and other intermediate generated data only to the system administrator. In other words, general users and administrators are not granted access to intermediate data. This clearly separates the entities that use the organization results for business purposes from those that can check the internal state for maintenance purposes. In particular, in this embodiment, access control is designed around the axis of "permission to view only to the system administrator," and general users and administrators are not allowed access to intermediate data beyond what is necessary for using the organization results.

[0102] The management terminal 11 shown in Figure 6 may be a terminal used by the system administrator. The management terminal 11 can access the internal state to the necessary extent via the control server 20, but it does not directly connect to the internal data store 40. In other words, even access by the management terminal 11 is not possible without going through the control of the control server 20 and the role-based access control unit 50. By limiting the access path even for internal access by the administrator in this way, it is possible to avoid direct exposure of the internal data store 40 while maintaining the necessary maintainability.

[0103] As shown in Figure 6, this embodiment ensures that text information, structured metadata, and feature representations are not transmitted to an external learning platform but are used solely for inference processing; the internal data store 40 does not have an external query acceptance interface and is accessible only from authenticated internal services; the transformation unit 22 and the organization unit 23 are activated only via an internal API 31 that is not publicly disclosed; and intermediate data is not displayed to general users or administrators. Therefore, this embodiment can function as an information processing platform that enhances security not only in terms of the organization result generation function but also in terms of isolated execution and access control.

[0104] Next, we will explain how to specifically apply the structure to an academic conference. In this embodiment, the target data is presentation information submitted to the academic conference. Presentation information may include presentation title, presentation abstract, presentation summary, submission category, keywords, etc. Note that the presentation abstract and presentation summary may be given as separate items or may be included in a single abstract field.

[0105] At academic conferences, it is necessary to organize a large amount of presentation information into consistent groups, making it easy to attend and ensuring consistency in content, by structuring it into a set of session units. Therefore, in this embodiment, the organizing unit 23 groups semantically similar presentations into sessions based on characteristic expressions generated from presentation titles, abstracts, summaries, submission categories, and keywords, thereby generating a set of session units. If necessary, the number or structure of each session may be adjusted while referring to the requirements for the number of presentations per session for each desired presentation format and other operational conditions.

[0106] In one example, the organization unit 23 may, in AI-driven mode, process all presentations submitted to the academic conference across the board and form groups of presentations based on the similarity between feature representations. In another example, in submission category-focused mode, processing may be performed for each existing submission category, or cluster generation or reorganization may be performed using the submission category as the basic unit. This allows for both an operation that broadly reflects semantic proximity and an operation that emphasizes consistency with existing classifications.

[0107] Furthermore, if necessary, proposed session titles may be generated for the generated sessions. These proposed session titles may be generated based on the representative themes, keywords, or structured metadata of the presentations belonging to each session, and may be used as supplementary information for the administrator to understand the session content and determine the display name.

[0108] Furthermore, the set of session units obtained by this embodiment may be passed on to subsequent program organization processes such as timetable arrangement, speaker duplication check, venue allocation, and other program organization processes. In other words, this embodiment is responsible for the stage of organizing semantically similar presentations into session units while securely handling presentation information, and subsequent processes such as timetable arrangement or venue resource allocation may be performed by other program organization processes.

[0109] Thus, when this embodiment is applied to an academic conference, it is possible to generate session-based collections using necessary attributes such as lecture title, lecture abstract, lecture summary, submission category, and keywords, while targeting lecture information that may include unpublished research content. Furthermore, by combining input control, inference-only use, deletion control, isolation execution, and access control, a secure organizational foundation suitable for handling transient data that is completed for each academic conference can be provided.

[0110] For example, if, in addition to lecture information, there are multiple lectures on crack detection in concrete structures, deterioration evaluation of tunnel linings, and anomaly detection in inspection records, the organization unit 23 may group the multiple lectures as candidates for the same session based on the similarity of these feature expressions and generate proposed session titles such as "Infrastructure Maintenance and Damage Detection." Subsequently, once the set of sessions is finalized or the academic conference has ended, the deletion control unit 25 may delete the feature expressions, structured metadata, and intermediate generated data used to generate the organization result. If it becomes necessary to retrieve the organization result again at a later date, instead of restoring it from deleted data, the analysis process may be performed again based on the original lecture information recorded in the original lecture information DB 41, and the organization result may be regenerated.

[0111] The above describes preferred embodiments of the invention disclosed herein applied to an information processing system for academic conferences. However, the invention disclosed herein is not limited to these embodiments. The target data is not limited to presentation information submitted to academic conferences, but may also include multiple document groups, proposal information groups, research theme information groups, review target information groups, and other descriptive information groups. In other words, the invention disclosed herein can be applied to various applications where it is necessary to securely handle text information contained in multiple target data and generate an organized result in which the multiple target data are organized into predetermined units. However, applications for academic conferences that handle presentation information that may include unpublished research content are one of the preferred applications of the invention disclosed herein.

[0112] Furthermore, the number, types, names, and granularity of attribute items or abstract categories used to generate structured metadata can be arbitrarily changed. For example, in one application, emphasis may be placed on subject area, research methods, target entities, and research objectives, while in another application, classification systems, review criteria, target users, or processing priorities may be used as semantic axes. Therefore, the attribute items or abstract categories are not fixed to the content or number exemplified in the implementation for academic conferences, but can be appropriately set according to the target field, operational purpose, or data characteristics.

[0113] Furthermore, the structured metadata generation method, feature representation generation method, clustering method, count condition adjustment method, and session title draft generation method can also be modified as appropriate. For example, structured metadata may be generated by inference using a language model, classification based on dictionaries or rules, or a combination thereof. Feature representations may be generated as numerical information or through compression processing. Grouping of presentations or target data sets may be implemented using density-based methods, hierarchical methods, affinity-based methods, constrained optimization methods, or other methods. Count condition adjustment or rearrangement processing may also be implemented using any method such as thresholding, iterative re-evaluation, or optimization calculations. For session title draft generation, any method such as representative word extraction, summary generation, category integration display, or other methods can be adopted.

[0114] The specific configuration of the internal data store 40 is not limited. The internal data store 40 can be any storage infrastructure capable of storing structured metadata, feature representations, and intermediate generated data, and may be implemented as a relational database, key-value storage, document storage, or other storage structure. Furthermore, the internal data store 40 may include a vector search infrastructure or a similar search data structure to assist in nearest neighbor search or similarity search of feature representations. For example, a vector index, a similarity search index, a nearest neighbor search graph structure, or other search support structures may be provided internally.

[0115] The implementation of the isolated execution environment 30 is not limited. The isolated execution environment 30 may be implemented as an on-premises environment, a virtual private cloud environment, a closed network configuration within a single cloud account, a container execution environment, or other managed execution environment. Furthermore, each block constituting the information processing system 1 may be aggregated on a single device or distributed across multiple devices. For example, the control server 20, the isolated execution environment 30, the internal data store 40, the original presentation information DB 41, and the organization result storage area 42 may be implemented logically separately on the same computing resource or distributed across different computing resources. Therefore, the invention of this disclosure is not limited to a specific infrastructure configuration or specific deployment form.

[0116] As described above, although the present invention has been explained using an embodiment for academic conferences as a preferred example, its technical concept can be broadly applied to various uses as a secure information processing platform that selects necessary text information from target data, converts said text information into structured metadata and feature representations, generates an organized result based on said feature representations, and further deletes internal data after processing is complete.

[0117] Each process shown in Figures 2 to 6 can be understood as an information processing method executed by one or more computers. For example, the information processing method may include the steps of acquiring target data, selecting necessary text attributes and excluding information that can directly identify an individual, analyzing the text information to generate structured metadata and feature representations, generating organization results based on the feature representations, saving or applying the organization results, deleting feature representations, structured metadata, and intermediate generated data in response to predetermined events, and re-executing based on the original data as necessary. The method may also include the steps of controlling non-transmission to an external learning platform, activating the conversion unit 22 and organization unit 23 within the isolated execution environment 30 via the internal API 31, and restricting the entities that can access the intermediate data using the role-based access control unit 50.

[0118] Furthermore, the present invention also includes a program for causing a computer to execute an information processing method. The program may cause the computer to function as an input control unit 21, a conversion unit 22, an organization unit 23, a learning control unit 24, a deletion control unit 25, and a role-based access control unit 50, or it may cause the computer to execute the corresponding steps. In other words, each function described as a system can be directly interpreted as a step in the method or a function in the program.

[0119] Furthermore, the program may be provided in the form of being recorded on a magnetic recording medium, optical recording medium, semiconductor memory or other computer-readable recording medium, or it may be provided in the form of being distributed or supplied via a communication line. The program, in whole or in part, may be executed by a single computer, or it may be executed by multiple computers in a divided manner. [Explanation of symbols]

[0120] 1. Information Processing System 10. User terminals 11 Management terminal 20 Control Servers 21 Input Control Unit 22 Conversion section 23 Organization Department 24 Learning Control Unit 25 Deletion Control Unit 30 Isolation Execution Environment 31 Internal API 40 Internal Datastore 41 Former Lecture Information Database 42 Organization result storage area 50 Role-Based Access Control Unit

Claims

1. An information processing system that analyzes text information contained in multiple target data and generates an organized result in which the multiple target data are organized into predetermined units, Conversion unit, The Organization Department, Deletion control unit and Equipped with, The conversion unit converts the text information into a feature representation that does not include a data structure that allows the original text to be directly restored. The aforementioned organization unit generates the organization result based on the feature representation, The deletion control unit deletes the feature expression when triggered by a predetermined event related to the completion of processing. Information processing system.

2. Further comprising an input control unit, The input control unit inputs only the text attributes necessary for generating the compilation result from the attributes included in the target data to the conversion unit, and excludes information that can directly identify an individual from the input to the conversion unit. The information processing system according to claim 1.

3. Further equipped with a learning control unit, The learning control unit controls the text information and feature representations not to be transmitted to an external learning platform, and to be used only for inference processing that generates the organization results, and not for additional training, fine-tuning, or permanent knowledge storage of the model. The information processing system according to claim 1.

4. The conversion unit is configured to generate the feature representation that does not retain specific lexical information, word order information, and syntactic structure by converting the information extracted from the text information into numerical information including category probability values ​​or weighted scores corresponding to a plurality of predefined attribute items or abstract categories. The information processing system according to claim 1.

5. The transformation unit is configured to generate the feature representation by applying at least one process selected from principal component analysis, random projection, and low-dimensional mapping processing to the numerical information corresponding to the attribute item or the abstract category, and by performing compression that involves information loss. The information processing system according to claim 4.

6. The aforementioned conversion unit is configured not to permanently store the intermediate analysis information generated by the analysis using the language model, but to discard it after the analysis is complete. The aforementioned intermediate analysis information includes at least one of the following: token sequence, attention weight, hidden state, detailed log of assignment to attribute item or abstract category, and individual application results of the mapping dictionary. The information processing system according to claim 4.

7. The system further includes an internal data store for storing the aforementioned feature representations and intermediate generated data generated during the organization process, The aforementioned internal data store is located in the internal network layer, isolated from the external network, does not have an interface for receiving external queries, and is accessible only from authenticated internal services. The information processing system according to claim 1.

8. Furthermore, equipped with a control server, The conversion unit and the organization unit are located within a managed, isolated execution environment, are activated in response to requests from the control server via an internal API that is not publicly exposed, and do not accept direct activation by general users or direct connections to the internal data store. The information processing system according to claim 7.

9. It further comprises a user interface unit and a role-based access control unit, The user interface unit receives registration of the target data, receives execution instructions from the organization unit, and displays the organization results, but does not display the feature representation and the intermediate generated data. The role-based access control unit classifies users into at least general users, administrators, and system administrators, and permits only the system administrators to view the feature representations and intermediate generated data. The information processing system according to claim 8.

10. The aforementioned data consists of presentation information submitted to academic conferences, which may include unpublished research content, and the data is complete for each academic conference. The information processing system according to claim 2.

11. The aforementioned text attributes include at least one of the following: lecture title, lecture summary, lecture abstract, submission category, and keywords. The information that can directly identify the aforementioned individual includes at least one of the author's name, affiliation information, contact information, session chair information, and presenter identification information. The information processing system according to claim 10.

12. The aforementioned organizational result is a collection of session units in the aforementioned academic conference, The deletion control unit deletes the feature representation, the structured metadata generated by the conversion unit, and the intermediate generated data generated by the organization unit, with respect to events related to the determination of the session unit set or the end of the academic conference as predetermined events. To retrieve the aforementioned organization results after such deletion, it is necessary to re-execute the analysis process based on the lecture information, and there is no function to restore the organization results from the feature representation, the structured metadata, and the intermediate generated data. The information processing system according to claim 10.

13. An information processing method performed by one or more computers, which analyzes text information contained in multiple target data and generates an organized result in which the multiple target data are organized into predetermined units, The conversion unit implemented by the one or more computers includes a conversion step of converting the text information into a feature representation that does not include a data structure that allows the original text to be directly restored, The organization unit implemented by the one or more computers includes an organization step of generating the organization result based on the feature representation, The deletion control unit implemented by one or more computers performs a deletion step in which the feature expression is deleted, triggered by a predetermined event related to the completion of processing. Information processing methods, including those mentioned above.

14. A program for causing a computer to execute the information processing method described in claim 13.