Data sovereignty-assured privacy-protected distributed genomic analysis system and method

The distributed genomic analysis system generates de-identified data within a client environment, ensuring secure transmission and local result recombination to address data leakage risks and maintain privacy, enabling accurate and customer-controlled genomic analysis.

KR102993420B1Active Publication Date: 2026-07-21EONEDIAGNOMICS CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
EONEDIAGNOMICS CO LTD
Filing Date
2025-08-01
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing genomic analysis platforms transmit raw genomic data and personal information to external environments, risking data leakage and violating privacy regulations, and lack customer control over data processing, while existing privacy protection technologies fail to separate necessary information from genomic data accurately.

Method used

A privacy-protected distributed genomic analysis system that generates de-identified data within a client environment, transmits it securely to an external server for analysis using encrypted APIs, and recombines results locally to maintain privacy and control data processing.

Benefits of technology

The system fundamentally blocks data leakage, ensures accurate analysis, and protects personal information by isolating raw data, enabling customer-led control and secure, real-time algorithm updates while preventing unauthorized use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112025087521026-PAT00008_ABST
    Figure 112025087521026-PAT00008_ABST
Patent Text Reader

Abstract

A genomic analysis system that performs analysis without external leakage of genomic data may include memory and a processor for storing instructions. When executed by the processor, the instructions can control the system to store raw genomic data in client-side storage, generate de-identified data from which personal information has been removed from the raw genomic data, transmit the de-identified data to an external server to request genetic analysis, receive analysis results from the external server, and generate a final analysis result in a client environment based on the received analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to the fields of bioinformatics and genomic data processing.

[0002] More specifically, the invention relates to a privacy-protected distributed genomic analysis system and a method of operation thereof, which performs genetic analysis by transmitting only de-identified data with personal information removed to an external server while securely storing raw genomic data and personal information in a client environment, and recombines the analysis results in the client environment to generate personalized results. Background Technology

[0004] The demand for genomic data processing is increasing due to the recent rapid proliferation of personal genomic analysis services. However, existing genomic analysis platforms primarily follow a structure where both raw genomic data and personal information are transmitted to external cloud environments to perform analysis. This approach carries risks of data leakage, violations of strengthened personal data protection regulations in various countries—such as the Personal Information Protection Act (PIPA), bioethics laws, GDPR, and HIPAA—and the leakage of sensitive genetic information.

[0005] In particular, genomic data contains sensitive personal information such as an individual's disease predisposition, drug responsiveness, and family history, which can cause irreversible and fatal damage once leaked. Furthermore, existing methods had a structural limitation in that customers could not proactively control the processing of their genomic data and had to rely entirely on external analysis companies.

[0006] Meanwhile, existing privacy protection technologies have primarily focused on data encryption or access control; however, due to the nature of genomic analysis, there is a problem in that it is difficult to obtain accurate analysis results while completely separating necessary personal information (e.g., age, gender, gestational age, BMI, phenotypic information, etc.) from genomic data.

[0007] Furthermore, there is a growing need for an API-based service model that can respond in real-time to the continuous development and updates of genetic analysis algorithms while preventing unauthorized copying or independent use of core algorithms. The problem to be solved

[0009] The present invention aims to provide a privacy-protected distributed genomics analysis system and method capable of performing accurate and reliable genetic analysis without transmitting raw genomic data and personal information outside the client environment.

[0010] The present invention aims to provide a system and method for generating de-identified data from raw genomic data that contains only the minimum information necessary for genetic analysis without being able to identify individuals, securely transmitting (via API) the data to an external server through multiple security layers in an encrypted form (e.g., "2.9321312.AT / 19.2123.AC") including chromosome number, position, and genetic information to perform analysis, and then recombining the results with the raw data in a client environment to generate personalized analysis results. Data processed on an external server has the characteristics of biologically fragmentary and context-removed garbage data, and therefore, even if stored, it may be impossible to identify individuals or extract meaningful information.

[0011] The present invention aims to protect the security of personal information by establishing a distributed server and an analysis module within a client environment so that all data processing and control are performed under client-led control, and by restricting communication with external servers to only the transmission of non-identifiable data and the reception of analysis results.

[0012] The present invention aims to provide an intellectual property protection service model that utilizes the latest genetic analysis algorithms and machine learning models from external servers through RESTful API-based communication, while preventing unauthorized copying or independent use of the algorithms and completely blocking API access upon termination of the service contract.

[0013] In addition, the privacy-protected anonymized data-based distributed genomic analysis system and method according to this document aims to protect rights at the national level from the perspective of data sovereignty and to prevent the overseas leakage of data for business purposes. means of solving the problem

[0015] A genomic analysis system that performs analysis without external leakage of the genomic data of this document may include memory for storing instructions and a processor. When executed by the processor, the instructions can control the system to store raw genomic data in a client environment, generate de-identified data from which personal information has been removed from the raw genomic data, transmit the de-identified data to an external server to request genetic analysis, receive analysis results from the external server, and generate a final analysis result in the client environment based on the received analysis results. Effects of the invention

[0017] The personal information protection-type anonymized data-based distributed genomic analysis system and method according to the present invention can fundamentally block the structural risk of data leakage by completely isolating and storing raw genomic data and personal information in a client environment.

[0018] In addition, the present invention can provide an optimized data processing effect that preserves key information necessary for analysis while making it impossible to identify individuals by applying de-identification technologies such as personal information identification and masking processing through pattern matching algorithms, prevention of backtracking through noise injection, and automatic generalization of rare patterns during the process of generating de-identified data.

[0019] The present invention has the effect of significantly improving data governance and reliability by securing autonomy in data processing through the establishment of distributed servers within a client environment, minimizing external dependency, and enabling customers to transparently control the entire analysis process.

[0020] The present invention can protect individual data through a customer approval-based selective data sharing system, while also contributing to the improvement of the accuracy of analysis algorithms utilizing anonymized aggregated data and the advancement of medical research.

[0021] The present invention enables multiple clients to obtain advanced analysis results utilizing collective intelligence while maintaining the security of their respective raw data through a collaborative analysis model, thereby simultaneously achieving personal information protection and improved analysis performance.

[0022] The personal information protection-type anonymized data-based distributed genomic analysis system and method according to this document can prevent data leakage overseas and guarantee national rights by controlling the system so that data does not leave the country, performs computations only through external Open APIs (RESTful APIs), and ensures that no data is stored externally. Brief explanation of the drawing

[0024] FIG. 1 is a diagram illustrating the overall structure of an artificial intelligence-based system according to one embodiment. FIG. 2 is a diagram illustrating the learning of a neural network according to one embodiment. FIG. 3 is a diagram illustrating the configuration of an artificial intelligence model according to one embodiment. Figure 4 is a block diagram showing the configuration of a personal information protection-type non-identification data-based distributed genomic analysis system according to one embodiment. FIG. 5 is a flowchart illustrating a distributed genomics analysis method based on personal information protection-type non-identification data according to one embodiment. FIG. 6 is a flowchart illustrating a distributed genomics analysis method based on personal information protection-type non-identification data according to one embodiment. FIG. 7 is a flowchart illustrating a distributed genomics analysis method based on personal information protection-type non-identification data according to one embodiment. FIG. 8 is a block diagram illustrating a data-protected genome analysis processing process within a client environment according to one embodiment. Specific details for implementing the invention

[0025] Hereinafter, embodiments are described in detail with reference to the attached drawings. However, various modifications may be made to the embodiments, and thus the scope of the patent application is not limited or restricted by these embodiments. It should be understood that all modifications, equivalents, and substitutions to the embodiments are included within the scope of the rights.

[0026] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Accordingly, the embodiments are not limited to the specific disclosed forms, and the scope of this specification includes modifications, equivalents, or substitutions that fall within the technical concept.

[0027] Terms such as "first" or "second" may be used to describe various components, but these terms should be interpreted solely for the purpose of distinguishing one component from another. For example, the first component may be named the second component, and similarly, the second component may be named the first component.

[0028] When it is stated that a component is "connected" to another component, it should be understood that it may be directly connected to or coupled with that other component, or that there may be other components in between.

[0029] The terms used in the embodiments are for illustrative purposes only and should not be interpreted as intended to be limiting. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0030] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the embodiments pertain. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0031] In addition, when describing with reference to the attached drawings, identical components are assigned the same reference numeral regardless of drawing symbols, and redundant descriptions thereof are omitted. In describing the embodiments, if it is determined that a detailed description of related prior art could unnecessarily obscure the essence of the embodiments, such detailed description is omitted.

[0032] The embodiments can be implemented in various forms of products such as personal computers, laptop computers, tablet computers, smartphones, televisions, smart home appliances, intelligent automobiles, kiosks, and wearable devices.

[0033] Artificial Intelligence (AI) systems are computer systems that implement human-level intelligence; unlike existing rule-based smart systems, they are systems in which machines learn and make decisions autonomously. As AI systems improve in recognition accuracy and gain a more accurate understanding of user preferences with continued use, existing rule-based smart systems are gradually being replaced by deep learning-based AI systems.

[0034] Artificial intelligence technology consists of machine learning and component technologies utilizing machine learning. Machine learning is an algorithmic technology that autonomously classifies and learns the characteristics of input data, while component technologies are technologies that mimic the cognitive and judgmental functions of the human brain by utilizing machine learning algorithms such as deep learning, and are comprised of technological fields such as linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control.

[0035] The various fields where artificial intelligence technology is applied are as follows. Linguistic understanding refers to technologies that recognize, apply, and process human language and text, including natural language processing, machine translation, dialogue systems, question answering, and speech recognition / synthesis. Visual understanding refers to technologies that perceive and process objects like human vision, including object recognition, object tracking, image search, people recognition, scene understanding, spatial understanding, and image enhancement. Inference and prediction refers to technologies that logically reason and predict by judging information, including knowledge / probability-based inference, optimization prediction, preference-based planning, and recommendation. Knowledge representation refers to technologies that automatically process human experiential information into knowledge data, including knowledge construction (data generation / classification) and knowledge management (data utilization). Motion control refers to technologies that control the autonomous driving of vehicles and the movement of robots, including motion control (navigation, collision, driving) and manipulation control (behavior control).

[0036] Generally, to apply machine learning algorithms to real-world situations, training is performed using a trial-and-error method due to the inherent characteristics of the fundamental methodologies. In particular, deep learning requires hundreds of thousands of iterations. Since it is impossible to execute this in a real physical external environment, training is instead performed through simulations that virtually recreate the actual physical environment on a computer.

[0037] In the present invention, Artificial Intelligence (AI) refers to a technology that imitates human learning ability, reasoning ability, and perceptual ability, and implements them on a computer, and may include concepts such as machine learning and symbolic logic. Machine Learning (ML) is an algorithmic technology that classifies or learns the characteristics of input data on its own. AI technology can analyze input data as a machine learning algorithm, learn from the results of the analysis, and make judgments or predictions based on the results of the learning. Furthermore, technologies that mimic the functions of the human brain, such as cognition and judgment, by utilizing machine learning algorithms can also be understood as falling within the category of AI. For example, technological fields such as linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control may be included.

[0038] Machine learning can refer to the process of training neural network models using experience in processing data. Through machine learning, computer software can improve its own data processing capabilities. A neural network model is constructed by modeling the correlations between data, and these correlations can be expressed by multiple parameters. A neural network model extracts and analyzes features from given data to derive correlations between them; machine learning can be defined as the process of optimizing the model's parameters by repeating this process. For example, a neural network model can learn the mapping (correlation) between inputs and outputs for data given as input-output pairs. Alternatively, even when only input data is provided, a neural network model can derive regularities between the given data and learn those relationships.

[0039] An artificial intelligence learning model or neural network model can be designed to implement the structure of the human brain on a computer and may include multiple network nodes that have weights and simulate neurons of a human neural network. The multiple network nodes may have interconnected relationships by simulating the synaptic activity of neurons, where neurons exchange signals through synapses. In an artificial intelligence learning model, multiple network nodes may be located in layers of different depths and exchange data according to convolutional connections. The artificial intelligence learning model may be, for example, an Artificial Neural Network (ANN) or a Convolutional Neural Network (CNN). As an embodiment, the artificial intelligence learning model may be machine learned according to methods such as supervised learning, unsupervised learning, and reinforcement learning. Machine learning algorithms for performing machine learning may include Decision Tree, Bayesian Network, Support Vector Machine, Artificial Neural Network, Ada-boost, Perceptron, Genetic Programming, and Clustering.

[0040] Among these, CNNs are a type of multilayer perceptron designed to use minimal preprocessing. CNNs consist of one or more convolutional layers and standard artificial neural network layers stacked on top, additionally utilizing weights and pooling layers. Thanks to this structure, CNNs can fully utilize two-dimensional input data. Compared to other deep learning architectures, CNNs demonstrate good performance in both image and audio fields. CNNs can also be trained using standard backpropagation. CNNs have the advantage of being easier to train than other feedforward artificial neural network techniques and using a small number of parameters.

[0041] Convolutional networks are neural networks comprising sets of nodes with bounded parameters. Many computer vision tasks have been significantly improved, driven by the increased size of available training data and the availability of computational power, combined with algorithmic advancements such as discriminative linear units and dropout training. In the case of massive datasets, such as those available for many tasks today, outfitting is not critical, and increasing the network size improves test accuracy. Optimal utilization of computing resources becomes a limiting factor. To address this, distributed, scalable implementations of deep neural networks can be employed.

[0043] FIG. 1 is a diagram illustrating the overall structure of an artificial intelligence-based system according to one embodiment.

[0044] As illustrated in FIG. 1, an artificial intelligence-based system (100) may include a plurality of user terminals (110-1, 110-n), a server (120), and a database (130). This system adopts a distributed architecture and operates based on a client-server model, with each component playing a unique role and contributing to increasing the efficiency and scalability of the entire system. According to one embodiment, the database (130) is depicted as being configured separately from the server (120), but is not limited thereto; depending on system design and operational efficiency, the database (130) may be integrated within the server (120). This integrated configuration has the advantage of improving data access speed and reducing system complexity. For example, the server (120) may include a plurality of artificial intelligence models and processing units for performing machine learning algorithms, and these may implement various types of deep learning and machine learning technologies (e.g., CNN, RNN, Transformer, reinforcement learning, etc.) to respond to user requests and provide intelligent services. According to one embodiment, a plurality of user terminals (110-1, 110-n), a server (120), and a database (130) can be connected to communicate with each other through a network (N), which enables real-time data exchange and smooth interaction between system components.

[0045] According to one embodiment, the network (N) may perform wireless or wired communication between a plurality of user terminals (110-1, 110-n), a server (120), a database (130), etc., and is located in the center of FIG. 1 and serves as a hub connecting all components. For example, the network may perform wireless communication according to methods such as 5G, LTE (Long-Term Evolution), LTE-A (LTE Advanced), CDMA (Code Division Multiple Access), WCDMA (Wideband CDMA), WiBro (Wireless BroadBand), WiFi (Wireless Fidelity), Bluetooth, NFC (Near Field Communication), GPS (Global Positioning System), or GNSS (Global Navigation Satellite System). The 5G network supports high-speed data transmission (up to 20Gbps), ultra-low latency (1ms or less), and large-scale connectivity, making it suitable for real-time processing of large-capacity AI models. For example, the network (N) may be configured to perform wired communication using methods such as USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface), RS-232 (Recommended Standard 232), Ethernet, fiber optic cable, or POTS (Plain Old Telephone Service). In particular, in data center environments, high-performance network technologies such as 400Gbps Ethernet or InfiniBand can be utilized for communication between large-scale AI computing clusters.

[0046] According to one embodiment, user terminals (110-1, 110-n) are various client devices that access the system and can be implemented in various forms such as smartphones, tablets, desktop computers, wearable devices, and IoT devices. These terminals are responsible for transmitting service requests through a user interface and receiving and displaying results processed by a server (120). Each terminal may also run a lightweight AI model locally, which can reduce network latency and enhance privacy protection. Direct connection or peer-to-peer (P2P) communication between terminals may be possible, as indicated by the dotted line, which is particularly useful in distributed learning or edge computing scenarios.

[0047] According to one embodiment, the server (120) is the central processing unit of the system, responsible for processing user requests, executing artificial intelligence models, and managing data. The server is equipped with a high-performance processor (CPU, GPU, TPU, NPU, etc.) for large-scale computation, large-capacity memory, and a stable operating system, and may be operated in a virtualized manner on a cloud infrastructure. The server performs training and inference functions for complex deep learning models and transmits the results to user terminals or stores them in a database. In addition, the stability and reliability of the system can be ensured through functions such as load balancing, error recovery, and security management.

[0048] According to one embodiment, a database (130) is a storage facility for storing and managing various data, and is depicted in a cylindrical shape in FIG. 1. Data stored in the database (130) is data acquired, processed, or used by at least one component of a plurality of user terminals (110-1, 110-n) or a server (120), and may include software (e.g., programs), user profiles, training datasets, model weights, log records, etc. Structured relational data may be stored in a SQL-based database (MySQL, PostgreSQL, etc.), and unstructured large-scale data may be stored in a NoSQL database (MongoDB, Cassandra, etc.) or a distributed file system (Hadoop HDFS, etc.). The database (130) may include volatile and / or non-volatile memory, where volatile memory (RAM) is used for caching or temporary data processing requiring fast data access, and non-volatile memory (SSD, HDD, tape drive, etc.) is used for permanent data storage. Databases also ensure the security and integrity of data through data encryption, access control, and backup and recovery mechanisms.

[0049] This artificial intelligence-based system (100) can be utilized in various application fields and can provide services such as user behavior analysis, recommendation systems, natural language processing, image and video recognition, and predictive analysis. In particular, the distributed architecture and scalable design of the system enable stable performance to be maintained even as the number of users and data volume increase.

[0051] FIG. 2 is a diagram illustrating the learning of a neural network according to one embodiment.

[0052] As illustrated in FIG. 2, the learning device can train a neural network (123) to classify review responses received from multiple user terminals (110-1, ���) by item. Additionally, the learning device can train a neural network (123) to extract user stay history from user movement path information. The neural network (123) is also called an artificial neural network and is a computational model designed inspired by the neural structure of the human brain, specialized in recognizing and learning complex patterns within data. According to one embodiment, the learning device may be a separate entity from the server (120), but it may also be implemented integrated into the same system, so it is not limited thereto.

[0053] According to one embodiment, the neural network (123) includes an input layer (121) into which training samples are input and an output layer (125) that outputs training outputs, and can be learned based on the difference between the training outputs and labels (i.e., actual correct data). Here, the labels are defined based on items corresponding to review responses (e.g., service quality, price satisfaction, cleanliness, etc.) and can be defined based on user dwell history (place visited, time spent, frequency of visit, etc.) corresponding to movement path information. The neural network (123) is connected as a group of multiple nodes and is defined by weights between the connected nodes and an activation function that activates the nodes. Various forms of the activation function may be used, such as sigmoid, hyperbolic tangent (tanh), and ReLU, which enables the network to learn non-linear patterns.

[0054] According to one embodiment, the learning device can train a neural network (123) using a Gradient Descent (GD) technique or a Stochastic Gradient Descent (SGD) technique. The GD technique is a method of updating weights all at once by calculating gradients based on the entire dataset, while the SGD technique is a method of increasing computational efficiency and the possibility of escaping a local optimum by updating weights more frequently using only a randomly selected portion of data (mini-batch). The learning device can use a loss function designed by the outputs and labels of the neural network. The loss function is an important element that quantifies the difference between the model's predicted value and the actual correct answer to suggest the direction of learning.

[0055] The learning device can calculate the training error using a predefined loss function. The loss function can be predefined with labels, outputs, and parameters as input variables, where the parameters can be set by weights within the neural network (123). For example, the loss function can be designed in the form of Mean Square Error (MSE), entropy, etc. MSE is calculated as the squared mean of the difference between the predicted value and the actual value and is mainly used for regression problems, while cross-entropy is a function that measures the difference between the predicted probability distribution and the actual distribution and is suitable for classification problems. In addition, loss functions that are less sensitive to outliers, such as Huber Loss, can be utilized, and various techniques or methods can be employed in the embodiments where the loss function is designed.

[0056] According to one embodiment, the learning device can identify weights that influence the training error using a backpropagation technique. Backpropagation is a process of calculating the degree to which each weight contributes to the final error while propagating the error calculated in the output layer toward the input layer, and is performed through differential calculations using the chain rule. Here, the weights are relationships between nodes within the neural network (123). The learning device may use an SGD technique using labels and outputs to optimize the weights found through the backpropagation technique. For example, the learning device may update the weights of a loss function defined based on labels, outputs, and weights using an SGD technique. This process operates by calculating the gradient (the derivative of the loss function with respect to the weights) and adjusting the weights in the direction of this gradient to gradually decrease the value of the loss function.

[0057] According to one embodiment, the learning device extracts first objects from a review response, obtains first labels which are items corresponding to the first objects, applies the first objects to a first neural network to generate first training outputs corresponding to the first objects, and can train the first neural network based on the first training outputs and the first labels. At this time, the first objects may include key keywords, sentence structures, sentiment expressions, etc. extracted from the review text, and thereby learn the ability to identify the core content of the review and classify it into the corresponding items.

[0058] According to one embodiment, the learning device extracts second objects from movement path information, obtains second labels which are user dwell history corresponding to the second objects, applies the second objects to a second neural network to generate second training outputs corresponding to the second objects, and can train the second neural network based on the second training outputs and the second labels. The second objects may consist of data such as the user's GPS coordinate sequence, movement speed, stopping point, and movement pattern, thereby learning the ability to determine where and for how long the user stayed.

[0059] According to one embodiment, a learning device can generate first training feature vectors based on constituent features (e.g., sentence structure, word frequency, part-of-speech distribution), positional features (e.g., location of important words, location of key sentences), and pattern features (e.g., repetitive expressions, usage patterns of specific phrases) of a review response. These feature vectors are generated through a process of converting raw text data into a numerical form that can be processed by a neural network, and various methods such as TF-IDF, Word2Vec, and BERT may be employed to extract features.

[0060] According to one embodiment, the learning device can generate second training feature vectors based on constituent features of movement path information (e.g., shape of the path, estimation of the means of movement), length features (e.g., total travel distance, travel time, distance between each point), and pattern features (e.g., repeated visit patterns, movement patterns by day of the week / time of the day). Since location data has continuous values ​​over time, time-series data processing techniques or spatial data analysis methodologies may be applied, and various methods may be employed to extract features.

[0061] According to one embodiment, the learning device can obtain training outputs by applying first training feature vectors to the neural network (123). This process is also called feedforward and is a process in which input data passes through each layer of the network to generate a final prediction value. The learning device can train the review item extraction algorithm of the neural network (123) based on the training outputs and first labels. The learning device can train the review item extraction algorithm of the neural network (123) by calculating training errors corresponding to the training outputs and optimizing the connection relationships of nodes within the neural network (123) to minimize the training errors. This process is repeated over several epochs, and in each epoch, the entire training dataset is processed and weights are updated. The server (120) can automatically extract and classify items from new review responses using the first neural network that has been trained.

[0062] According to one embodiment, the learning device can obtain training outputs by applying second training feature vectors to the neural network (123). The learning device can train the user stay history acquisition algorithm of the neural network (123) based on the training outputs and second labels. The learning device can train the user stay history acquisition algorithm of the neural network (123) by calculating training errors corresponding to the training outputs and optimizing the connection relationships of nodes within the neural network (123) to minimize the training errors. Since movement path data has temporal continuity, a recurrent neural network (RNN) structure such as LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit) can be applied, which enables effective learning of the patterns of the sequence data. The server (120) can obtain user stay history from movement path information using the second neural network after training is complete, and the information thus obtained can be utilized in various application fields such as improving location-based services, building customized recommendation systems, and analyzing traffic patterns.

[0064] FIG. 3 is a diagram illustrating the configuration of an artificial intelligence model according to one embodiment.

[0065] An artificial intelligence model according to one embodiment may include an input layer, a hidden layer, and an output layer. This multi-layer structure is effective for learning complex patterns and solving non-linear problems, and is utilized as a core structure, particularly in deep learning models. The input layer is a layer related to input values ​​input into the artificial intelligence model and serves to receive data coming from the outside. As illustrated in FIG. 3, the input layer consists of three nodes, and each node may represent a specific feature or attribute of the data. The number of nodes in the input layer may vary depending on the dimension or characteristics of the data to be processed, and may be determined by the number of pixels in the case of image data, the dimension of the word vector in the case of text data, etc.

[0066] According to one embodiment, a feature map can be output by performing a MAC (multiply-accumulate) operation and an activation operation on the input values ​​in the hidden layer. In FIG. 3, the hidden layer consists of four nodes, and these nodes are connected to all nodes of the input layer to form a fully connected structure. The MAC operation may be an operation that multiplies the input values ​​by their corresponding weights and sums the multiplied values, and this acts as the basic unit of information processing in a neural network.

[0067] According to one embodiment, the activation operation may be an operation that inputs the result of a MAC operation into an activation function and outputs a result value. The activation function is a key element that enables a model to learn complex patterns by introducing non-linearity into the result of a linear operation. The activation function may be of various types. For example, the activation function may include a sigmoid function (which outputs a value between 0 and 1 and is useful for expressing probability), a tangent function (which outputs a value between -1 and 1 and is suitable for balanced data), a ReLU function (which outputs 0 for negative inputs and outputs it as is for positive inputs, mitigating the vanishing gradient problem), a Leaky ReLU function (which applies a linear function with a small gradient to negative inputs), a MaxOut function (which selects the maximum value among several linear functions), and / or an ELU function (which is an exponential linear unit function that provides a smooth saturation curve for negative inputs), but there are no limitations on the types thereof. Recently, new activation functions such as Swish and Mish are also being studied and are contributing to the improvement of model performance.

[0068] According to one embodiment, the hidden layer may be composed of at least one layer, and the more layers there are, the more complex patterns can be learned, resulting in a deep learning structure. For example, when the hidden layer is composed of a first hidden layer and a second hidden layer, the first hidden layer performs MAC operations and activation operations based on the input value of the input layer to output a feature map, and the feature map, which is the result value from the first hidden layer, can become the input value for the second hidden layer. The feature map is an intermediate representation representing features extracted from input data; in image processing, it can represent visual features such as edges, textures, and shapes, and in text processing, it can represent semantic patterns or syntactic structures. The second hidden layer can perform MAC operations and activation operations based on the feature map, which is the result value of the first hidden layer, and through this hierarchical processing, increasingly abstract and high-dimensional features can be extracted.

[0069] According to one embodiment, the output layer may be a layer associated with the result of an operation performed in the hidden layer, and in FIG. 3, it consists of three nodes. The number of nodes in the output layer varies depending on the type of problem to be solved; for binary classification problems, it may be one node, for multi-class classification problems, it may be equal to the number of classes, for regression problems, it may be determined by the dimension of the value to be predicted, and for sequence generation problems, it may be determined by the length of the output sequence. In the output layer, a special activation function such as a softmax function or a sigmoid function is generally used to generate a final predicted value or a probability distribution.

[0070] In one embodiment, the learning model learns syllable (character) patterns that are frequently combined and used in a given corpus to automatically learn the boundaries of compound words and named entities, integrates object information from a first UI source with object information rendered in a browser to create a learning object information file, and uses the learning object information file to generate training data for training a deep learning network. In this process, natural language processing (NLP) technology is utilized, and architectures such as word embedding, recurrent neural networks (RNN), and Transformers may be applied. Additionally, the learning model receives data from various domains of a support system, standardizes the data from the various domains into an integrated format based on at least one standardization method corresponding to each of the various domains, learns and infers data from a specific domain, determines information to be transmitted for standardization from the data of the specific domain, and performs post-processing on the data from the various domains.

[0071] According to one embodiment, the first UI source includes an XML file, and the training object information file includes an input JSON file for feature learning and an output JSON file that serves as label data during training. This structured data format allows the model to process it easily and is particularly useful for analyzing UI elements of web-based applications. The output JSON file includes a file containing HTML DOM Tree information implemented in compliance with web standards, and the various domains include at least one of a RAN (radio access network), a transport, or a core. A RAN refers to a wireless access network where a mobile device connects to a cellular network; a transport refers to a network layer responsible for data transmission; and a core refers to the central part of a network responsible for central processing and routing. Post-processing may include a correlation function, which plays an important role in analyzing relationships between data collected from various domains and discovering meaningful patterns. Through correlation analysis, the model can identify the causes of anomalies or performance degradation occurring in complex systems and improve prediction accuracy.

[0073] Figure 4 is a block diagram showing the configuration of a personal information protection-type non-identification data-based distributed genomic analysis system according to one embodiment.

[0074] A system (400) according to one embodiment may include a processor (420) and memory (430), and some of the illustrated components may be omitted or substituted. Although the system diagram briefly depicts only the core components, the actual implementation may include various additional components such as an input / output controller, a system bus, a graphics processing unit (GPU), a network interface card (NIC), and a storage device controller. A system (400) according to one embodiment may be a server or a terminal; if implemented as a server, it may be operated in a virtualized environment as part of a cloud infrastructure or implemented as a physical hardware server, and if implemented as a terminal, it may be implemented in various forms such as a desktop computer, a laptop, a tablet, a smartphone, an embedded system, or an IoT device. According to one embodiment, the processor (420) is a component capable of performing operations or data processing regarding the control and / or communication of each component of the system (400), and may be composed of one or more processors. In modern processor architectures, multi-core designs are common, which improve parallel processing performance by integrating multiple independent processor cores within a single chip. For example, configurations such as dual-core, quad-core, and octa-core are possible, and each core has independent cache memory (L1, L2), and some cache (L3) can be shared among cores.

[0075] The memory (430) can store information related to the method described above or a program in which the method described above is implemented. The memory subsystem is designed with a hierarchical structure and is composed of several levels according to access speed and capacity. The memory (430) may be volatile memory or non-volatile memory. Volatile memory is memory in which stored data is lost when the power is turned off, and is mainly implemented in the form of DRAM (Dynamic RAM) or SRAM (Static RAM). DRAM stores each bit in the form of charge in a capacitor and requires periodic refreshing, whereas SRAM uses flip-flop circuits to maintain data, so it does not require refreshing but has higher cost and power consumption due to the use of more transistors. The memory (430) can store various file data, and the stored file data can be updated according to the operation of the processor (420). The memory management unit (MMU) is responsible for converting virtual memory addresses to physical memory addresses, managing page tables, and maintaining memory protection and cache consistency.

[0076] According to one embodiment, the processor (420) can execute a program and control the device (400). The processor can operate through basic pipeline stages of instruction fetch, decode, execute, and write-back. The code of the program executed by the processor (420) can be stored in memory (430). In addition to the application, the operating system, device driver, middleware, system services, etc. are loaded into memory and executed, and these provide basic system functions such as hardware resource management, process scheduling, interrupt handling, and file system management. The operations of the processor (420) can be performed by loading instructions stored in memory (430). In this process, the program counter (PC) points to the address of the next instruction to be executed, the instruction register (IR) stores the instruction currently being executed, and various general-purpose registers store operation data and intermediate results. The system (400) can be connected to an external device (e.g., a personal computer or a network) through an input / output device (not shown in the drawing) and exchange data. The input / output system can operate on an interrupt-based or polling-based basis and can transfer data directly between memory and the input / output device without the intervention of a processor through a Direct Memory Access (DMA) controller.

[0077] According to one embodiment, data exchange between memory (430) and processor (420) is performed via a system bus, which consists of an address bus, a data bus, and a control bus. The address bus specifies the memory location that the processor intends to access, the data bus transmits actual data, and the control bus transmits control information such as read / write signals. The system (400) may be operated under an operating system that supports advanced features such as multitasking, virtual memory, memory protection, and separation of privileges, and if a real-time operating system (RTOS) is used, it may satisfy deterministic response times and strict time constraints. For artificial intelligence applications, the memory (430) stores a neural network model, weight parameters, a training dataset, intermediate calculation results, etc., and the processor (420) may perform inference and learning by executing forward propagation and backpropagation algorithms. The system (400) can ensure high availability and fault tolerance by applying technologies such as clustering, load balancing, and failover mechanisms for scalability, and can implement technologies such as data encryption, access control, secure booting, and memory protection for security.

[0079] FIG. 5 is a flowchart illustrating a distributed genomics analysis method based on personal information protection-type non-identification data according to one embodiment.

[0080] The operations described through FIG. 5 can be implemented based on instructions that can be stored in a computer recording medium or memory (e.g., memory (430) of FIG. 4). The order of each operation of FIG. 5 may be changed, some operations may be omitted, and some operations may be performed simultaneously.

[0081] In operation 510, the system (400) can store raw genomic data in a client environment (e.g., client storage) and generate de-identified data. The system (400) can securely store raw genomic data in the format of FASTQ, BAM (Binary Alignment / Map), or VCF (Variant Calling Format) in a secure storage within the client environment. FASTQ is a raw data file containing base sequences and quality scores, BAM is a binary file that stores the results after aligning the FASTQ file to a reference sequence, and VCF may refer to a file format containing variant information. The system (400) can generate de-identified data by removing information that can identify an individual from the stored raw genomic data, thereby protecting the data and fundamentally blocking the risk of external leakage. The system (400) can verify the integrity of the raw data through a distributed processing module and extract only the minimum information necessary for analysis to comply with personal information protection regulations.

[0082] In operation 520, the system (400) can request genetic analysis by transmitting non-identified data to an external server. The system (400) can encrypt the generated non-identified data and then securely transmit it to an external analysis server via an HTTPS-based secure channel. The system (400) can access the latest genetic analysis algorithm service of the external server using a RESTful API interface and can continuously update the algorithm. The system (400) can control the transmission process so that raw genome data and personal information are completely isolated, thereby improving data security.

[0083] In operation 530, the system (400) can receive analysis results from an external server and generate a final analysis result. The system (400) can receive the analysis results processed by the external server in an encrypted form. The analysis results may be in the form of a code that cannot identify an individual. The system (400) can decrypt the received analysis results within a client environment and combine them with raw genome data and personal information to generate personalized genome analysis results. Through this analysis process, the system (400) can provide an accurate and meaningful analysis report and can provide high-quality analysis services without external exposure of personal information throughout the entire process.

[0084] According to one embodiment, the system (400) can store raw genome data in a client environment, generate non-identifiable data from which personal information has been removed from the raw genome data, transmit the non-identifiable data to an external server to request gene analysis, receive an analysis result from the external server, and control the generation of a final analysis result in the client environment based on the received analysis result.

[0085] According to one embodiment, the system (400) can protect data and fundamentally block the risk of external leakage by storing raw genomic data in a client environment. The system (400) can store genomic data in various formats such as FASTQ, BAM, and VCF on client-side local storage, a private cloud, or an on-premises server. Through this storage method, the system (400) can satisfy the personal information protection regulations and data localization requirements of each country. The system (400) can ensure the transparency and reliability of data governance by configuring the client to have complete control over access rights to the stored data. The storage format and location of the genomic data are merely examples and are not limited thereto, and may vary depending on the client's security policy and infrastructure environment.

[0086] According to one embodiment, the system (400) can achieve an optimal balance between privacy protection and analysis performance by generating non-identifiable data from which personal information has been removed from raw genomic data. The system (400) can utilize a distributed processing module to systematically detect and remove not only direct identifiers included in the raw data but also information that can indirectly identify an individual, while preserving variant information and sequence data essential for genetic analysis.

[0087] According to one embodiment, the system (400) can effectively defend against personal information leakage and re-identification attacks while preserving the usefulness of data to the maximum extent by applying advanced de-identification techniques, including a pattern matching algorithm, statistical noise injection, and data generalization techniques, in a multi-layered manner. A pattern matching algorithm may refer to a technique that identifies and extracts strings or data having a specific rule or pattern defined in advance within the data using a regular expression or a machine learning model. Through this, the system (400) can prioritize the detection of direct personal identification information, such as names, contact information, and resident registration numbers, and process them for masking or deletion.

[0088] According to one embodiment, statistical noise injection may refer to a technique of adding intentional random values ​​(noise) to data within a range that does not impair the statistical significance of the analysis results. For example, by injecting fine noise into aggregated information such as the frequency of specific gene variants or average age, it is possible to prevent differencing attacks—where an attacker queries analysis results multiple times to trace back the information of a specific individual—and to satisfy the principle of Differential Privacy. Data generalization techniques may refer to a technique that reduces the specificity of information by converting the precise values ​​of data into a higher category. The types and application methods of de-identification techniques are merely examples and are not limited thereto; they may vary depending on the characteristics of the data and the purpose of the analysis.

[0089] According to one embodiment, the system (400) can protect personal information while using the latest analysis technology by transmitting non-identifiable data to an external server to request genetic analysis. The system (400) can interact with an external analysis server through a standardized communication protocol based on RESTful API, thereby enabling real-time utilization of continuously updated genetic analysis algorithms and machine learning models. The system (400) can ensure data integrity and confidentiality during transmission by applying multiple security layers, and can restrict access to the analysis service to only authenticated clients through access control via an API gateway. The communication method and security protocol with the external server are merely examples and are not limited thereto, and may vary depending on the network environment and security requirements.

[0090] According to one embodiment, the system (400) can obtain results of precise genetic analysis utilizing high-performance computing resources by receiving analysis results from an external server. The system (400) can receive the analysis results in the form of an encrypted code composed of a combination of alphabets and numbers. Since the encrypted code itself cannot be understood in terms of meaning, the risk of information leakage during the transmission process can be minimized. The system (400) can decrypt the received results by releasing the security layer applied during transmission in reverse order, and can ensure the reliability of the analysis results through data integrity verification. The format and encryption method of the analysis results are merely examples and are not limited thereto, and may vary depending on the characteristics of the analysis service and security requirements.

[0091] According to one embodiment, the system (400) can complete a personalized genomic analysis service by generating a final analysis result in a client environment based on the received analysis result. The system (400) can securely combine the decoded analysis result with raw genomic data and personal information stored in the client environment through a result generation module, thereby providing comprehensive health information regarding an individual's genetic characteristics, disease risk, drug responsiveness, nutrient metabolism, etc. The system (400) can ensure an accurate connection between the analysis result and personal data by utilizing a mapping table and a unique identification number. The system (400) can completely prevent external exposure of personal information by controlling the identification process so that it takes place only within the client environment. The components and provision form of the final analysis result are merely examples and are not limited thereto, and may vary depending on the individual's requirements and the purpose of the analysis.

[0093] FIG. 6 is a flowchart illustrating a distributed genomics analysis method based on personal information protection-type non-identification data according to one embodiment.

[0094] The operations described through FIG. 6 can be implemented based on instructions that can be stored in a computer recording medium or memory (e.g., memory (430) of FIG. 4). The order of each operation of FIG. 6 may be changed, some operations may be omitted, and some operations may be performed simultaneously.

[0095] In operation 610, the system (400) can identify personal information using a pattern matching algorithm through a distributed processing module and perform masking processing. The system (400) can automatically detect personal information from raw genomic data, including age, gender, gestational age, BMI, and phenotypic information including eye color, skin color, and hair loss. The types of personal information are merely examples and are not limited to this, and may vary depending on the settings. The system (400) can utilize advanced pattern recognition technology to comprehensively identify not only direct identifiers but also information that can indirectly identify an individual. The system (400) can apply masking to the identified personal information so that the individual cannot be identified externally, while preserving key information necessary for genetic analysis.

[0096] In operation 620, the system (400) can assign a random unique number and perform multi-encryption processing based on an algorithm. The system (400) can assign a unique identification number to each non-identified data through a random number generator in a client environment, thereby enabling accurate matching of the analysis results with the raw data later. The system (400) can perform primary encryption of the non-identified data using a verified public encryption algorithm including AES (Advanced Encryption Standard) or RSA (Rivest-Shamir-Adleman).

[0097] In operation 630, the system (400) can communicate with an external server through a RESTful API-based HTTPS / SSL secure channel. The system (400) can form a secondary encrypted secure channel by utilizing the HTTPS protocol and an SSL (Secure Sockets Layer) security certificate, thereby enhancing data security at the transport layer. The system (400) can establish a standardized interface with the external server in JSON or XML data format through web-based communication in the form of a RESTful API. The system (400) can prevent man-in-the-middle attacks or eavesdropping during the data transmission process through a secure channel with TLS (Transport Layer Security) encryption applied, and can guarantee the integrity of the communication protocol.

[0098] In operation 640, the system (400) can decrypt the analysis results and combine them with raw data to generate personalized results. The system (400) can receive analysis results in the form of unidentifiable codes composed of combinations of alphabets and numbers from an external server. The system (400) can decrypt the received analysis results by applying a multi-encryption method in reverse order. The system (400) can accurately match the corresponding raw genomic data and personal information by referring to the decrypted analysis result code and a mapping table stored in the client environment. Through this combination process, the system (400) can generate comprehensive and personalized genomic analysis results that include at least one of an individual's genetic characteristics, disease risk, or drug responsiveness, and can completely block the risk of external exposure of personal information throughout the entire process.

[0099] According to one embodiment, the system (400) can generate non-identified data by using a pattern matching algorithm to identify personal information including age, gender, gestational age, BMI (Body Mass Index), and phenotypic information including eye color, skin color, and hair loss status in raw genome data through a distributed processing module installed in a client environment, and by removing it through masking processing. In addition, the non-identified data is processed to include only the minimum information necessary for genetic analysis without being able to identify an individual from the outside, and a random unique number generated in the client environment is assigned to the non-identified data to manage it so that the analysis results can be matched with the raw data later, and the non-identified data is first encrypted using an encryption algorithm and can be transmitted to an external server by forming a secure channel that is secondarily encrypted through an SSL (Secure Sockets Layer) security certificate based on the HTTPS (Hypertext Transfer Protocol Secure) protocol. In addition, it communicates with an external server through web-based communication in the form of a RESTful API (Representational State Transfer Application Programming Interface), and when receiving analysis result data in the form of unidentifiable code from the external server, it decrypts it by applying the same multi-encryption method in reverse order; it generates personalized genomic analysis results by combining the received analysis result code with raw genomic data and personal information stored in the client environment within the client environment, and controls are made so that the raw genomic data and personal information are not transmitted outside the client environment during the entire analysis process.

[0100] According to one embodiment, the system (400) can ensure the robustness of privacy protection by systematically identifying and removing personal information from raw genomic data through a distributed processing module installed within a client environment. The system (400) can comprehensively detect not only direct personal information such as age, gender, gestational age, and BMI, but also phenotypic information including eye color, skin color, and hair loss by utilizing a pattern matching algorithm. The system (400) can accurately identify even hidden personal information by combining advanced pattern recognition techniques including regular expressions, machine learning-based classifiers, and statistical outlier detection, and can safely remove or transform the identified information by applying various processing techniques including masking, hashing, or tokenization.

[0101] According to one embodiment, the system (400) can accurately identify and remove even hidden personal information by organically combining advanced pattern recognition techniques including regular expressions, machine learning-based classifiers, and statistical anomaly detection.

[0102] A regular expression may refer to a technique for detecting a specific type of string based on a predefined rule. For example, the system (400) can predefine a unique pattern of a resident registration number, such as '6-digit number - 7-digit number', or a unique combination format of a patient ID used in a specific hospital, as a regular expression, and can quickly identify information matching the pattern within the raw data.

[0103] Additionally, a machine learning-based classifier may refer to an artificial intelligence model that learns the characteristics of personal information from a large amount of data and identifies personal information within new data. The system (400) can effectively classify and detect even unstructured personal information by understanding the context of text recorded in the form of comments or memos through natural language processing technology and probabilistically determining whether a combination of specific words or expressions is likely to identify an individual.

[0104] Statistical outlier detection may refer to a technique that analyzes the distribution of the entire data to identify rare values ​​that deviate significantly from general patterns. Rare variants or unique combinations of variants that appear only in specific individuals in genomic data can serve as indirect identifiers capable of identifying individuals. Therefore, the system (400) can prevent the risk of re-identification in advance by automatically detecting statistically significant outliers through comparison with a population genomic database. The types of personal information and pattern matching techniques are merely examples and are not limited thereto, and may vary depending on the characteristics of the dataset and privacy requirements.

[0105] According to one embodiment, the system (400) can achieve an optimal balance between privacy protection and analysis accuracy by processing non-identifiable data so that it contains only the minimum information necessary for genetic analysis, while ensuring that an individual cannot be identified externally. The system (400) can mathematically control the risk of individual re-identification by applying privacy preservation techniques such as differential privacy. Differential privacy may refer to a technique that reduces the risk of personal information leakage while utilizing data analysis or query results. The system (400) can protect personal information by adding random noise to the dataset, thereby minimizing the impact of a specific individual's information on the results and making it difficult to identify a specific individual's information solely from the analysis results.

[0106] According to one embodiment, the system (400) can simultaneously preserve information critical to genetic analysis, including variation frequency, association analysis, and population statistics, in order to maintain statistical significance. The system (400) can quantitatively evaluate and optimize data usability and privacy protection levels using information-theoretical metrics, and can dynamically adjust the scope and level of granularity of necessary information according to the purpose of analysis. Privacy preservation techniques and information minimization methods are merely examples and are not limited thereto, and may vary depending on analysis requirements and the regulatory environment.

[0107] According to one embodiment, the system (400) can assign a random unique number generated in the client environment to non-identifying data to ensure an accurate match between the analysis result and the raw data, while controlling the connection relationship so that it cannot be identified externally. The system (400) can generate a unique identifier that is unpredictable and has an extremely low probability of collision by utilizing a cryptographically secure random number generator, and such identifier can store connection information with the raw data only in a mapping table within the client environment. The system (400) can establish a system that is meaningless externally but allows for accurate connections internally through a hash function, UUID generation, or symmetric key-based tokenization technique. The system (400) can also form an additional security layer by encrypting the mapping table itself. The unique number generation method and mapping management technique are merely examples and are not limited thereto, and may vary depending on the security policy and system architecture.

[0108] According to one embodiment, the system (400) can ensure the confidentiality of the data content itself and comply with international standard security criteria by first encrypting non-identified data using an encryption algorithm. The system (400) can defend against cryptanalysis attacks by applying verified public encryption algorithms including AES-256, RSA-2048, and ChaCha20. The system (400) can systematically manage the generation, distribution, circulation, and disposal of encryption keys through a key management system, and can enhance the security of the keys by utilizing a hardware security module or a key derivation function. The types of encryption algorithms and key management methods are merely examples and are not limited thereto, and may vary depending on security requirements and performance considerations.

[0109] According to one embodiment, the system (400) can implement additional security at the transport layer and prevent man-in-the-middle attacks by forming a secondary encrypted secure channel through an SSL security certificate based on the HTTPS protocol. The system (400) can ensure confidentiality by applying a modern transport security protocol such as TLS 1.3, and can verify the identity of the server and confirm the identity of the communication counterpart through a digital certificate issued by a certification authority. The system (400) can generate an independent encryption key per session and simultaneously guarantee data integrity by combining key exchange algorithms including ECDHE (Elliptic Curve Diffie-Hellman Ephemeral) and DHE (Diffie-Hellman Ephemeral) with the AEAD (Authenticated Encryption with Associated Data) encryption mode, and can block protocol downgrade attacks through certificate pinning or HSTS policies. ECDHE and DHE may refer to algorithms used to securely exchange keys in TLS / SSL communication. AEAD can refer to an encryption mode that encrypts data while simultaneously ensuring data integrity. HTTPS security settings and certificate management methods are merely examples and are not limited to this; they may vary depending on the network environment and security policies.

[0110] According to one embodiment, the system (400) can ensure compatibility with various analysis services by configuring a standardized and scalable interface with an external server through web-based communication in the form of a RESTful API. The system (400) can reduce coupling between systems and improve scalability through JSON or XML-based data exchange or stateless communication methods. The system (400) can guarantee the stability and security of the service through API version management, request restrictions, and authentication token-based access control, and can efficiently respond to large-scale data analysis requests by utilizing asynchronous processing or webhooks. The API communication method and data format are merely examples and are not limited thereto, and may vary depending on the specifications and performance requirements of the external service.

[0111] According to one embodiment, when the system (400) receives analysis result data in the form of unidentifiable code from an external server, it can safely restore the analysis result while maintaining data integrity and confidentiality by decrypting it by applying the previously applied multiple encryption method in reverse order. The system (400) can sequentially perform transport layer decryption through TLS session termination and application layer decryption using a symmetric or asymmetric key, and can verify whether the data has been tampered with by verifying a message authentication code or digital signature at each stage. The decryption method and integrity verification technique are merely examples and are not limited thereto, and may vary depending on the encryption configuration and security requirements.

[0112] According to one embodiment, the system (400) can provide personalized health information under complete privacy protection by generating personalized genomic analysis results by combining the received analysis result code, raw genomic data stored in the client environment, and personal information within the client environment. The system (400) can accurately link the anonymized analysis results with the individual's raw data by referring to a mapping table. The system (400) can adjust genetic risk levels individually and generate personalized health recommendations by considering phenotypic information including at least one of the individual's age, gender, or body mass index. The system (400) can present personalized preventive medical guidelines by comprehensively analyzing various genetic characteristics, including disease susceptibility, drug responsiveness, nutrient metabolism, and exercise responsiveness, and can provide complex genetic information in an easy-to-understand form through visualization tools or dashboards. The components of the analysis results and the method of personalization are merely examples and are not limited thereto, and may vary depending on the purpose of the analysis and the individual's health condition.

[0113] According to one embodiment, the system (400) can protect data and naturally comply with the personal information protection regulations of each country by controlling the transmission of raw genomic data and personal information outside the client environment during the entire analysis process. The system (400) can also prevent the unintended leakage of sensitive information at the application level through data classification and labeling. The system (400) can record all data access and processing history in a traceable manner through audit logs. Data leakage prevention techniques and monitoring methods are merely examples and are not limited thereto, and may vary depending on the organization's security policies and regulatory requirements.

[0114] According to one embodiment, the system (400) installs a distributed processing module for genomic analysis on a distributed server built within a client environment and generates analysis results on the distributed server to fundamentally block the path for raw genomic data and source data containing personal information to be exposed outside the client environment, and is configured so that all data processing and control are performed within the client environment, and communication with an external server is restricted only to the transmission of non-identifiable data and the reception of analysis results, thereby allowing the client to proactively control the entire analysis process.

[0115] According to one embodiment, the system (400) can protect data and minimize external dependencies by installing a distributed processing module for genomic analysis on a distributed server built within the client environment. The system (400) can deploy analysis functions, including genomic data preprocessing, quality verification, and mutation detection, to the client's local environment by utilizing container-based virtualization technology or a microservices architecture, thereby enabling advanced analysis of sensitive data without it crossing the client boundary. The system (400) can efficiently utilize the client's hardware resources through a distributed computing framework. The distributed server construction method and analysis module deployment technique are merely examples and are not limited thereto, and may vary depending on the client's infrastructure environment and performance requirements.

[0116] According to one embodiment, the system (400) can generate analysis results through a distributed server to fundamentally block the path through which source data containing raw genomic data and personal information is exposed outside the client environment. The system (400) can physically or logically isolate sensitive data processing areas by applying network segmentation technology that divides the network into several small logical networks (segments or subnets).

[0117] According to one embodiment, the system (400) can detect and block unintended copying or movement of source data in real time through a data classification policy and an automated monitoring system, and can ensure the transparency and integrity of the data processing history through a blockchain-based audit trail or an immutable log system. Data isolation techniques and leakage prevention methods are merely examples and are not limited thereto, and may vary depending on security policies and regulatory requirements.

[0118] According to one embodiment, the system (400) can ensure transparency in data governance and maximize the client's autonomous operational capabilities by configuring all data processing and control to take place within the client environment. The system (400) can support the client in directly setting and monitoring the entire process from data collection to the generation of analysis results through a workflow management system and a task scheduler. The system (400) can transparently visualize the progress of analysis and the system status through a real-time dashboard and a notification system. The system (400) can autonomously manage data retention periods, access rights, and processing priorities according to the client's internal regulations by utilizing policy-based automation tools, and can continuously check the status of regulatory compliance by automatically generating audit reports and compliance checklists. The internal control system and governance tools are merely examples and are not limited thereto, and may vary depending on the organization's operational policies and management framework.

[0119] According to one embodiment, the system (400) may allow only minimal external connections by restricting communication with an external server to only the transmission of non-identifiable data and the reception of analysis results. The system (400) may allow only communication with pre-approved external analysis services through a whitelist-based firewall policy, and may block abnormal data transmission attempts even through allowed communication channels. The system (400) may prevent attempts at large-scale data leakage through API rate limiting and request size limits, and may monitor and warn of communication patterns or data transmission volumes that differ from normal in real time through an abnormal behavior detection system. Communication restriction techniques and anomaly detection methods are merely examples and are not limited thereto, and may vary depending on network security policies and threat environments.

[0120] According to one embodiment, the system (400) can protect data and reduce dependency on external vendors by allowing the client to proactively control the entire analysis process. The system (400) can subdivide analysis control rights according to the authority structure within the organization through role-based access control and a multi-stage approval workflow, and can be configured so that the client can control key decisions such as whether to use external services, the scope of data sharing, and how to utilize analysis results. The process control method and authority management system are merely examples and are not limited thereto, and may vary depending on the organization's decision-making structure and responsibility system.

[0122] According to one embodiment, the system (400) separates and stores raw genomic data and personal information including age, gender, gestational age, BMI, and phenotypic information on a distributed server within a client environment, and processes the raw genomic data through a distributed processing module installed on the distributed server to convert it into a numerical value for final determination calculation that cannot identify an individual and can be interpreted only within the internal system, thereby generating non-identifiable data. Furthermore, the non-identifiable data is processed in a form that prevents the inference or restoration of original genetic information and is configured to have no meaning in itself. A random unique number generated through a random number generation algorithm in the client environment is assigned to each non-identifiable data to maintain a mapping table within the client only that can subsequently match the raw data with analysis results received from an external server. Additionally, during all communication processes with the external server, the raw genomic data and personal information can be controlled to be processed only within an isolated internal processing area so that they cannot access the network transmission layer.

[0123] According to one embodiment, the system (400) can fundamentally block the risk of data fusion attacks and strengthen the robustness of privacy protection by storing raw genomic data and personal information separately on distributed servers within a client environment. The system (400) can utilize database sharding techniques or a distributed file system to place genomic sequence information and phenotypic metadata in physically separated repositories, and can be configured so that complete personal information is not exposed even if one repository is compromised by applying different encryption keys and access rights. The system (400) can manage identifier information and sensitive attributes in separate tables through data schema partitioning and normalization techniques, and can implement the principle of least privilege by restrictively sharing only the key information required for join operations. The data separation storage method and access control technique are merely examples and are not limited thereto, and may vary depending on the database architecture and security requirements.

[0124] According to one embodiment, the system (400) can ensure security by converting raw genomic data into numerical values ​​that cannot identify individuals and can be interpreted only within the internal system through a distributed processing module installed on a distributed server. The system (400) can convert high-dimensional genomic data into low-dimensional feature vectors by applying dimensionality reduction techniques such as principal component analysis, singular value decomposition, or autoencoders. These numerical values ​​cannot be inversely converted back to the original sequence information, while preserving statistical patterns and correlations. The system (400) can implement a one-way conversion in which computational results can be obtained but input data cannot be restored by utilizing hash-based feature extraction or homomorphic encryption techniques, and can be configured to convert even the same data into different numerical values ​​by using unique conversion parameters for each client. The numerical value conversion technique and feature extraction method are merely examples and are not limited thereto, and may vary depending on the purpose of analysis and security requirements.

[0125] According to one embodiment, the system (400) can achieve complete privacy protection from an information-theoretic perspective by configuring the non-identifiable data to be processed in a form from which the original genetic information cannot be inferred or restored, so that it has no meaning in itself. The system (400) can generate unpredictable outputs for the same input by combining an irreversible hash function and a salt value, and can be designed so that the original data cannot be recovered from the outside. The system (400) can mathematically verify that the transformed data is statistically independent of the original through information entropy measurement, and can quantitatively evaluate the possibility of backtracking through mutual information analysis and control it below a security threshold. The original unrecoverable processing method and security verification technique are merely examples and are not limited thereto, and may vary depending on the cryptographic security model and the level of privacy guarantee.

[0126] According to one embodiment, the system (400) can protect privacy while maintaining connectivity by assigning a random unique number generated through a random number generation algorithm in a client environment to each non-identifying data, and maintaining a mapping table only within the client that can match the analysis results received from an external server with the raw data. The system (400) can generate an identifier that is unpredictable and free from statistical bias by combining a cryptographically secure pseudo-random number generator and a hardware entropy source, and can detect and avoid identifier collisions by utilizing a hash table. The system (400) can be configured to maintain the mapping table itself only in memory or store it in a temporary encrypted storage so that it is automatically erased upon system restart. The unique number generation method and mapping table management technique are merely examples and are not limited thereto, and may vary depending on system security policies and performance requirements.

[0127] According to one embodiment, the system (400) can implement data protection at the physical level by controlling that raw genomic data and personal information are processed only in an isolated internal processing area so that they cannot access the network transmission layer during all communication processes with an external server. The system (400) can place sensitive data processing processes in a separate execution environment and perform data operations in an isolated memory space that does not have a direct interface with the network stack. The system (400) can fundamentally block sensitive information from reaching the network output buffer or transmission queue through data flow tracking and information flow control mechanisms, and can configure an additional physical isolation layer by utilizing a hardware security module or a trusted execution environment. The network isolation technique and the method of configuring the internal processing area are merely examples and are not limited thereto, and may vary depending on the hardware platform and security architecture.

[0129] FIG. 7 is a flowchart illustrating a distributed genomics analysis method based on personal information protection-type non-identification data according to one embodiment.

[0130] The operations described through FIG. 7 can be implemented based on instructions that can be stored in a computer recording medium or memory (e.g., memory (430) of FIG. 4). The order of each operation of FIG. 7 may be changed, some operations may be omitted, and some operations may be performed simultaneously.

[0131] In operation 710, the system (400) can transmit non-identified data by sequentially applying three-dimensional multiple security layers. The system (400) can fundamentally block the possibility of personal identification through a first security process that de-identifies the data itself, and then encrypt the data content through a second security process using an encryption algorithm. The system (400) can add security at the transport layer by applying a third security process through a security channel based on an SSL certificate of the HTTPS protocol. Through this layered security structure, the system (400) can respond to different security threats at each stage and maximize the robustness of data security by establishing multiple lines of defense.

[0132] In operation 720, the system (400) can control the external server to perform genetic analysis operations on non-identifiable data. The system (400) can process complex operations such as mutation analysis, disease prediction, and drug responsiveness analysis by utilizing the high-performance computing resources and the latest genetic analysis algorithms of the external server. The system (400) can restrict the external server to perform analysis only on non-identifiable data and can configure the system architecture so that it cannot access raw genomic data or personal information. The system (400) can continuously update the analysis algorithm through an API-based service model and prevent unauthorized copying or independent use of the core algorithm.

[0133] In operation 730, the system (400) receives the analysis results in the form of unidentifiable code and can decrypt them by applying the same multiple security layers in reverse order. The system (400) can receive the result data returned after the completion of genetic analysis from an external server in the form of an encrypted code composed of a combination of alphabets and numbers, and this code itself can be designed so that no meaning can be understood. The system (400) can decrypt the third security processing (SSL / TLS), second security processing (public algorithm encryption), and first security processing (de-identification) that were applied during the transmission process in reverse order. Through this stepwise decryption process, the system (400) can verify the integrity of the analysis results and restore the data in a form that can be interpreted only in a client environment.

[0134] In operation 740, the system (400) can generate a personalized genomic analysis report in a result generation module. The system (400) can securely combine the decrypted analysis result code with raw genomic data and personal information stored within the client through a dedicated result generation module installed within the client environment. The system (400) can generate a detailed report containing health management guidelines and preventive recommendations specialized for each individual by comprehensively analyzing the individual's genetic variations, disease risk, drug responsiveness, and nutrient metabolism characteristics. The system (400) can control the entire analysis process to proceed autonomously within the client environment, and the external server can protect the privacy of the data and the individual by limiting its role to merely providing computation services for non-identifiable data.

[0136] According to one embodiment, the system (400) may configure multiple security layers by sequentially applying a first security processing that de-identifies the data itself when transmitting non-identified data to an external server, a second security processing that encrypts the de-identified data using an encryption algorithm, and a third security processing through a security channel based on an SSL certificate of the HTTPS protocol. Additionally, when receiving analysis result data returned after the completion of genetic analysis by the external server, the system may receive the data in the form of an unidentifiable code composed of a combination of alphabets and numbers, decrypt it by applying the same multiple security layers in reverse order, and a result generation module installed within the client environment may combine the received code with raw genome data and personal information stored within the client to generate a personalized genome analysis result report. The system architecture may be configured such that the entire analysis process proceeds independently within the client environment, and the external server is limited to performing only analysis operations on the non-identified data.

[0137] According to one embodiment, the system (400) can overcome the limitations of a single security technique by sequentially applying three multiple security layers and can establish a deep defense system capable of maintaining overall security even in the event of a failure of defense at each layer. The system (400) can provide comprehensive protection against complex threats such as brute force attacks, man-in-the-middle attacks, and protocol vulnerabilities by designing each security layer to respond to different attack vectors, and can be configured to maintain the security of other layers even if one layer is compromised by using security parameters independent of each layer. The system (400) can immediately detect data tampering or loss by adding checksums or integrity tags at each security processing step, and can ensure service continuity through automatic retransmission or the use of alternative paths in the event of failure. The configuration method of multiple security layers and the failure response strategy are merely examples and are not limited thereto, and may vary depending on the threat model and security requirements.

[0138] According to one embodiment, the system (400) can maximize the semantic ambiguity of the result data itself and minimize the risk of information exposure during transmission by receiving the analysis result from an external server in the form of an unidentifiable code composed of a combination of alphabets and numbers. The system (400) can convert the numerical result into a text form. The system (400) can be designed so that the original meaning cannot be understood without the decryption key or conversion table of the corresponding client. The system (400) can conceal the actual result size through code length normalization and padding techniques, and can prevent pattern analysis attacks through the insertion of dummy data or shuffling. The code conversion method and concealment technique are merely examples and are not limited thereto, and may vary depending on the data characteristics and security requirements.

[0139] According to one embodiment, the system (400) can ensure data integrity and authentication and detect attempts by attackers to manipulate data through a verification mechanism at each stage when decrypting by applying multiple security layers in reverse order. The system (400) can sequentially perform HMAC-based message authentication, digital signature verification, checksum comparison, etc., by applying different verification algorithms for each layer, and can immediately stop decryption and generate a security warning if verification fails at any stage. The system (400) can block old data or duplicate transmissions through timestamp verification and retransmission attack prevention techniques, and can prevent the insertion of malicious payloads by verifying the format and range of the result data even after successful decryption. The reverse decryption verification method and attack detection technique are merely examples and are not limited thereto, and may vary depending on security policies and threat environments.

[0140] According to one embodiment, the system (400) can provide precision medical information that comprehensively considers an individual's genetic characteristics and health status by generating a personalized genomic analysis result report through a result generation module installed within a client environment. The system (400) can express complex genetic information in an easy-to-understand chart or graph through a visualization library. The result generation method and personalization technique are merely examples and are not limited thereto, and may vary according to user requirements and medical guidelines.

[0141] According to one embodiment, the system (400) can clearly separate responsibilities by configuring a system architecture in which the entire analysis process is conducted independently within the client environment, and the external server is limited to performing analysis operations on non-identifiable data. The system (400) can track and record all activities of the external server through audit logs and behavior monitoring, and can immediately block the connection and send a warning to the client upon detection of abnormal behavior. The method of configuring the system architecture and the technique of separating roles are merely examples and are not limited thereto, and may vary depending on service requirements and security policies.

[0143] According to one embodiment, the system (400) may select a structure that utilizes the analysis service of an external server via an API call method without installing an analysis engine on the client in order to respond in real time to version updates of the gene analysis algorithm provided by the external server. Additionally, the system may be configured to allow access only through an API gateway so that the analysis algorithm and machine learning model of the external server are not copied or stored in the client environment in the form of binary files or source code, thereby preventing unauthorized copying and independent use of the algorithm. The system (400) may call and use the latest version of the gene analysis service provided by the external server through an API endpoint, but when the service contract is terminated, it may invalidate the API authentication token so that the client can no longer access the core algorithm, and control the system to process mutation analysis and disease prediction operations on non-identified gene sequence data with personal information removed by utilizing the computing resources of the external server.

[0144] According to one embodiment, the system (400) can minimize the burden on the client-side system by utilizing an analysis service from an external server via an API call method without installing an analysis engine on the client. The system (400) can delegate complex genome analysis logic to an external specialized service, thereby allowing the client to focus solely on core data protection functions and reduce the burden of maintaining analysis algorithms. The system (400) can ensure stable integration even when external services are updated through version compatibility management and API schema verification, and can provide an alternative analysis path in the event of a service failure through a fallback mechanism. The API-based service usage method and integration management technique are merely examples and are not limited thereto, and may vary depending on the service architecture and operational requirements.

[0145] According to one embodiment, the system (400) can protect intellectual property rights and prevent technology leakage by configuring the system so that analysis algorithms and machine learning models from an external server are accessed only through an API gateway, thereby preventing them from being copied or stored in the client environment. The system (400) restricts the client to receiving only the execution results of the algorithm and can fundamentally block direct access to model parameters or training data. The system (400) can prevent model extraction attacks through code obfuscation and server-side execution environment isolation, and can protect the service provider's revenue model. The intellectual property protection method and access control technique are merely examples and are not limited thereto, and may vary depending on license policies and security requirements.

[0146] According to one embodiment, the system (400) can establish clear contract-based service boundaries and prevent unauthorized use by invalidating the API authentication token when the service contract is terminated, thereby preventing the client from accessing the core algorithm any further. The system (400) can grant the client time-limited and scope-limited access rights and manage the continuous authentication status through token expiration or renewal. The system (400) can implement immediate token invalidation through a centralized token management server and blacklist management, and the client side can also completely remove traces of access through token deletion and cache clearing. The token-based authentication method and invalidation mechanism are merely examples and are not limited thereto, and may vary depending on the authentication policy and security protocol.

[0147] According to one embodiment, the system (400) can obtain high-performance analysis results while reducing the burden of hardware investment on the client by utilizing computing resources of an external server to process mutation analysis and disease prediction calculations on non-identifiable gene sequence data. The system (400) can dynamically allocate resources according to the analysis workload through elastic computing and automatic scaling, and can efficiently perform parallel analysis of large-scale genomic data by utilizing GPU accelerated computation or distributed processing clusters. The system (400) can efficiently process analysis requests from multiple clients through task queue management and priority scheduling, and can support economical service usage through resource usage monitoring and cost optimization. The method of utilizing computing resources and performance optimization techniques are merely examples and are not limited thereto, and may vary depending on the cloud platform and analysis requirements.

[0149] According to one embodiment, the system (400) performs a procedure to obtain explicit approval for a genome analysis summary result generated in a client environment through a customer's electronic signature or clicking an approval button via a user interface, and can display only the approved summary result in a state where it can be extracted from a database and transmitted to an external server. The summary result may consist of aggregated data including the frequency distribution of individual genetic variations, statistical average values ​​of disease risk, and categorized classification results of drug responsiveness, but with unique identifiers capable of identifying individuals removed. The external server stores the received summary result as an anonymized dataset and applies license conditions that restrict the scope of use to machine learning training data for improving the accuracy of analysis algorithms or for statistical research purposes only, and can control the client to modify matters including the selection of data items to be shared, the setting of the sharing period, and restrictions on the purpose of use through a data sharing settings interface.

[0150] According to one embodiment, the system (400) can guarantee an individual's right to self-determination regarding data and comply with the consent-based data processing principles required by the GDPR or personal data protection laws by performing an explicit consent procedure through a user interface. The system (400) can utilize electronic signatures or biometric authentication and induce informed consent by clearly notifying the purpose and scope of data usage during the consent process. The system (400) can ensure the transparency and traceability of the consent process by recording the consent history and timestamps on a blockchain or audit log. The explicit consent method and consent management technique are merely examples and are not limited thereto, and may vary depending on legal requirements and user interface design.

[0151] According to one embodiment, the system (400) can prevent indiscriminate data sharing and control that only information permitted by the user is transmitted externally by extracting only approved summary results from the database and displaying them in a state where they can be transmitted to an external server. The system (400) can automatically exclude unapproved data from the transmission target and immediately remove it from the transmission queue upon withdrawal of approval, and can block unauthorized access to unapproved data through permission-based access control. The system (400) can automatically exclude even approved data from the sharing target after a set period has passed through data lifecycle management and automatic expiration policies, and can immediately notify the user of changes in the data sharing status. The approval-based data management method and state control technique are merely examples and are not limited thereto, and may vary depending on the data governance policy and system architecture.

[0152] According to one embodiment, the system (400) can achieve a balance between the protection of personal information and the value of research utilization by processing the summary results as aggregated data in which the frequency distribution of individual genetic variations, the statistical mean value of disease risk, and the categorized classification results of drug responsiveness are composed, while removing personal identifiers. The system (400) can convert individual genetic information into meaningful patterns and trends by utilizing statistical aggregation techniques and data mining algorithms. The system (400) can mathematically manage the risk of re-identification of individuals during the aggregation process by applying differential privacy techniques. Differential privacy may refer to a technique that reduces the risk of personal information leakage while utilizing data analysis or query results. The system (400) can improve data quality by filtering out noise or outliers and can ensure compatibility with other research data by using a standardized medical terminology system. The method of configuring aggregated data and privacy preservation techniques are merely examples and are not limited thereto, and may vary depending on the research purpose and medical requirements.

[0153] According to one embodiment, the system (400) can ensure the ethical use of data by applying license conditions to store summary results received by an external server as an anonymized dataset and utilize them only for machine learning training or statistical research purposes. The system (400) can automatically monitor data usage conditions through smart contracts and block access in case of violation, and can detect use for purposes other than authorized purposes through usage log analysis and audit trails. The system (400) can track unauthorized redistribution or tampering by applying data watermarking or fingerprinting technology, and can support automatic notification and legal action in case of violation by linking with legally binding data usage agreements. License management methods and usage restriction techniques are merely examples and are not limited thereto, and may vary depending on legal frameworks and ethical guidelines.

[0154] According to one embodiment, the system (400) can protect personal data and support dynamic privacy management by allowing a client to finely control sharing conditions through a data sharing setting interface. The system (400) can support users in setting differentiated sharing policies by data item, purpose, and period by providing an intuitive dashboard and granular permission setting options. The system (400) can support rollback to previous settings through change history management. The data sharing control interface and policy management techniques are merely examples and are not limited thereto, and may vary depending on user experience requirements and system complexity.

[0156] According to one embodiment, the system (400) manages the non-identifiable data by assigning a random unique number generated using a random number generator in a client environment to the non-identifiable data through a mapping table so that the analysis results and the raw data can be linked later, and the non-identifiable data can be first encrypted using an encryption algorithm including AES (Advanced Encryption Standard) or RSA (Rivest-Shamir-Adleman). In addition, it can form a secure channel with TLS (Transport Layer Security) encryption applied through an SSL security certificate based on the HTTPS protocol and transmit it to an external server, communicate with the external server in JSON or XML data format through web-based communication in the form of a RESTful API, and when receiving analysis result data in the form of an unidentifiable code composed of a combination of alphabets and numbers from the external server, it can decrypt it by applying the same multiple encryption method in reverse order. Furthermore, by referring to the received analysis result code and the mapping table of the client environment, it can generate a personalized genomic analysis result by combining the corresponding raw genomic data and personal information, and during the entire analysis process, it can control the transmission of raw genomic data and personal information through a network interface so that they are not transmitted outside the client environment.

[0157] According to one embodiment, the system (400) can minimize security risks caused by the preservation of long-term data connection information by utilizing an in-memory database or volatile storage when managing unique numbers through a mapping table, thereby ensuring that the data is automatically erased upon system restart. The system (400) can independently manage the mapping relationship between unique numbers and raw data on a session-by-session basis and can completely block subsequent connection tracking by deleting the mapping information immediately after analysis is completed. The system (400) can protect against insider attacks or memory dump attacks by applying encryption and access control to the mapping table itself, and can prevent single-point failure through distributed storage. The mapping table security management method and lifecycle control technique are merely examples and are not limited thereto, and may vary depending on the system security policy and memory management strategy.

[0158] According to one embodiment, the system (400) can simultaneously utilize the speed of symmetric key encryption and the key management advantages of asymmetric key encryption by applying a hybrid encryption method that combines AES and RSA. The system (400) can optimize processing performance by securely exchanging AES session keys through RSA and high-speed encrypting actual data with AES, and can periodically update the encryption key. The hybrid encryption implementation method and key management strategy are merely examples and are not limited thereto, and may vary depending on performance requirements and security policies.

[0159] According to one embodiment, the system (400) can ensure the safety of past communication content and prevent certificate spoofing attacks even when the session key is exposed by utilizing full forward secrecy and certificate fixation techniques when forming a secure channel with TLS encryption applied.

[0160] According to one embodiment, the system (400) can detect data format errors and tampering in advance and guarantee the integrity of API calls by utilizing JSON schema verification and XML digital signatures in RESTful API communication. The system (400) can support the efficient transmission of large-capacity genomic data by utilizing compression algorithms and chunking transmission. The system (400) can maintain connectivity with existing clients even during service upgrades through API version management and backward compatibility guarantees, and can prevent service abuse through rate limiting and quota management. API communication optimization techniques and data verification methods are merely examples and are not limited thereto, and may vary depending on service characteristics and performance requirements.

[0161] According to one embodiment, the system (400) can selectively allow only permitted communication patterns using whitelist-based firewall rules and can perform real-time threat response through automatic isolation and warning generation upon detection of abnormal traffic. The system (400) can logically separate sensitive data processing areas by utilizing network virtualization and overlay technology. Network security control techniques and traffic blocking methods are merely examples and are not limited thereto, and may vary depending on the network architecture and security policy.

[0163] According to one embodiment, the system (400) prevents backtracking of the original data by injecting noise into the variant information of the raw genome data during the process of generating non-identifying data, and blocks the risk of re-identification by automatically generalizing or removing the relevant part when a rare pattern that can identify an individual is detected among the gene sequences, and controls data usage so that it is traceable by tracking repetitive analysis requests for the same data and restricting additional analysis if the number of times exceeds a set limit, and by recording the generation and transmission history of all non-identifying data transmitted externally in an immutable manner.

[0164] According to one embodiment, the system (400) can mathematically eliminate the possibility of backtracking individual data points while maintaining statistical significance by injecting noise into the variation information of raw genomic data. The system (400) can add randomness to the original variation frequency through a differential privacy mechanism and can optimize the balance between analysis accuracy and privacy protection by dynamically adjusting the noise intensity. The system (400) can generate alternative data that is similar to the actual variation pattern but cannot be linked to a specific individual by combining it with a synthetic data generation technique, and can restore the original data when necessary by encrypting and storing the noise injection history. The noise injection technique and privacy budget management method are merely examples and are not limited thereto, and may vary depending on data characteristics and privacy requirements.

[0165] According to one embodiment, the system (400) can prevent the risk of individual re-identification due to unique genetic characteristics in advance by automatically detecting rare patterns in gene sequences and generalizing or removing them. The system (400) can identify statistically rare variations or combinations through comparative analysis with a population genetics database and can automatically recognize unique individual patterns by utilizing a machine learning-based outlier detection algorithm.

[0166] According to one embodiment, the system (400) can prevent multiple attacks or attempts at information leakage by tracking repetitive analysis requests for the same data and restricting additional analysis when a set threshold is exceeded. The system (400) can determine identity through data fingerprinting technology and hash-based identification. The system (400) can apply independent restriction policies by user, IP, and session.

[0167] According to one embodiment, the system (400) can ensure complete audit trail and accountability by recording the creation and transmission history of all non-identifiable data transmitted externally in an immutable manner. The system (400) can cryptographically guarantee the integrity of log data by utilizing blockchain technology or a hash chain, and can ensure the reliability of the recording time by linking with a timestamp service. The system (400) can record the entire transformation process from the original data to the final transmission in detail through data lineage tracking and version control, and can provide necessary evidence upon request by regulatory or auditing bodies. The immutable recording method and audit trail system are merely examples and are not limited thereto, and may vary depending on legal requirements and compliance standards.

[0169] FIG. 8 is a block diagram illustrating a data-protected genome analysis processing process within a client environment according to one embodiment.

[0170] The system (400) can protect sensitive genomic data by clearly separating the client storage (810) area from the external environment (830). The client storage (810) is a physically or logically isolated security area that can safely store all sensitive information, including personal information such as raw genomic data (812), customer data (820), and final analysis results (824), and through this isolation structure, the risk of external intrusion or data leakage can be fundamentally blocked. The configuration method of the client storage is merely an example and is not limited to this, and may vary depending on security requirements.

[0171] According to one embodiment, the system (400) may store raw genome data (812) in various standard formats such as FASTQ, SAM, BAM, and VCF, and such data may include an individual's whole gene sequence information, single nucleotide polymorphisms (SNPs), structural variations (SVs), copy number variations (CNVs), etc. The system (400) may store metadata associated with the raw genome data, such as an individual's age (e.g., 35 years), gender (e.g., female), gestational age (e.g., 28 weeks), BMI (e.g., 23.5), and phenotypic information (e.g., brown eyes, black hair, no hair loss), thereby supporting correlation analysis between genetic variations and phenotypes. The format of the raw genome data and the configuration of the metadata are merely examples and are not limited thereto, and may vary depending on the purpose of analysis and data standards.

[0172] According to one embodiment, the system (400) can generate de-identified data (816) by systematically identifying and removing personal information from raw genome data through a distributed processing module (814). The distributed processing module (814) can comprehensively detect not only direct identifiers (e.g., name, resident registration number) but also indirect identifiers (e.g., combinations of rare genetic variants, variant patterns specific to a particular region) by utilizing advanced technologies such as pattern matching algorithms, machine learning-based classifiers, and statistical outlier detection, and can apply various processing techniques such as masking (e.g., replacing actual values ​​with special characters), hashing (e.g., one-way transformation via SHA-256), and generalization (e.g., processing into age ranges instead of exact ages). The configuration techniques of the distributed processing module are merely examples and are not limited thereto, and may vary depending on the de-identification requirements.

[0173] According to one embodiment, the system (400) can transmit non-identified data (816) to an analysis module (832) in an external environment (830) to perform advanced genetic analysis. The non-identified data (816) is information converted into numerical values ​​or code forms (e.g., "A1B2C3", "0.7542", "TYPE_X") that cannot identify an individual, and although its meaning cannot be understood externally, information such as disease risk, drug responsiveness, and nutrient metabolism characteristics can be derived through a specialized analysis algorithm. The system (400) can apply encryption such as AES-256 or RSA-2048 and form a secure channel through the HTTPS / TLS 1.3 protocol to enhance data protection during transmission. The form of the non-identified data and the encryption method are merely examples and are not limited thereto, and may vary depending on the security policy.

[0174] According to one embodiment, the system (400) may receive the results analyzed in an external environment (830) as result data (818) in the form of an encrypted code, and this result data may be designed to consist of a combination of alphabets and numbers so that its meaning cannot be understood from the outside. The system (400) may decrypt the received result data (818) within a client environment and, by referring to a pre-generated mapping table, securely combine it with the individual's raw genome data (812) and customer data (820), thereby generating a final analysis result (824) that includes health information and recommendations specialized for each individual. The code form of the result data and the decryption process are merely examples and are not limited thereto, and may vary depending on data security requirements.

[0175] According to one embodiment, the system (400) can provide personalized health management guidelines tailored to an individual's genetic characteristics when generating a final analysis result (824) through a result processing server (822). For example, if a specific individual has a genotype of slow caffeine metabolism, specific recommendations (e.g., "Please limit your daily caffeine intake to 200 mg or less and avoid consuming caffeine after 2 p.m.") are examples only and are not limited thereto, and may vary depending on the individual's genetic characteristics and health condition.

[0176] According to one embodiment, the system (400) controls access so that sensitive information, including raw genome data (812), customer data (820), and final analysis results (824), is not transmitted outside the boundaries of the client storage (810) during the entire analysis process, and controls so that only non-identifiable data (816) and result data (818) are exchanged with the external environment (830) in a limited manner. Through this architecture, the system (400) can naturally satisfy strict regulatory requirements such as the Personal Information Protection Act, the Bioethics Act, and GDPR. Network isolation and access control methods are merely examples and are not limited thereto, and may vary depending on the security architecture.

[0178] According to one embodiment, the system (400) can perform collaborative analysis while multiple client environments possess their own raw data, by having an external server distribute an analysis model to each client, and each client can train the model with its own non-identifiable data and then encrypt only the training results and transmit them to the external server. Additionally, the external server can control the system to ensure the reliability of the analysis results while protecting privacy by applying a verification mechanism that integrates the received encrypted training results to improve the overall analysis model, but prevents the raw data or non-identifiable data of each client from being transmitted outside the respective client environment, evaluates the data quality and analysis contribution of each client to exclude data with low reliability from the overall analysis, and verifies the accuracy of intermediate results generated during the collaborative analysis process without disclosing the actual data content.

[0179] According to one embodiment, the system (400) can analyze data collectively while maintaining individual data security through the Federated Learning (FL) paradigm by performing collaborative analysis while multiple client environments hold their own raw data. The Federated Learning paradigm may refer to a technology in which multiple local clients and a single central server cooperate to learn a global model in a decentralized data environment. The system (400) can coordinate multiple clients to participate in learning simultaneously by utilizing a distributed machine learning framework and model parameter synchronization technology. The learning implementation method and collaboration coordination technique are merely examples and are not limited thereto, and may vary depending on the network environment and the scale of participating clients.

[0180] According to one embodiment, the system (400) can protect data privacy by having an external server distribute an analysis model to each client and collect only the learning results. The system (400) can be configured to encrypt and transmit only model weight or gradient information, and to prevent the central server from directly checking the learning content of individual clients. The system (400) can hide the influence of individual data points by applying differential privacy to the learning results. The model distribution and aggregation method and privacy preservation technique are merely examples and are not limited thereto, and may vary depending on model complexity and security requirements.

[0181] According to one embodiment, the system (400) can prevent a decrease in learning quality by evaluating the data quality and analysis contribution of each client and excluding data with low reliability. The system (400) can quantify the data reliability of each client through statistical outlier detection and measurement of model performance contribution.

[0182] According to one embodiment, the system (400) can simultaneously guarantee transparency and privacy by applying a proof mechanism that verifies the accuracy of intermediate results during the collaborative analysis process but does not disclose the actual data content. The system (400) can independently verify each step of the calculation process by utilizing checksum-based integrity verification and a Merkle tree structure, and can guarantee the immutability of the verification results by linking with smart contracts or blockchain technology. The zero-knowledge verification mechanism and distributed verification method are merely examples and are not limited thereto, and may vary depending on cryptographic protocols and verification requirements.

[0184] According to one embodiment, the system (400) can generate non-identifiable data by applying differential privacy technology. Beyond simply masking or removing personal identification information, the system (400) can inject mathematically calculated noise into the results of data queries to make it impossible to determine whether a specific individual is included in the dataset. According to one embodiment, the system (400) can encrypt the data using homomorphic encryption technology when transmitting the non-identifiable data to an external server. Homomorphic encryption may refer to encryption technology that allows the data to be processed (e.g., statistical analysis, model training) on ​​an external server while remaining in an encrypted state. In this case, the external server can perform analysis without decrypting the ciphertext and then return the encrypted result value to the client. The system (400) can obtain the final analysis result by decrypting this encrypted result value only within the client environment. Through this, the external server cannot know the contents of the non-identifiable data or the original data at all, so the risk of information leakage can be theoretically eliminated throughout the entire data processing process. The types and scope of application of homomorphic encryption algorithms are merely examples and are not limited thereto; they may vary depending on the computational complexity required for analysis and the performance requirements of the system.

[0185] According to one embodiment, the system (400) can perform distributed analysis in which multiple clients cooperate by implementing federated learning. In this structure, an external server only distributes a central analysis model to each client, and each client can train the model in a local environment without exposing its raw data to the outside. Subsequently, only local learning results, such as parameters of the trained model (e.g., weights) rather than the raw data, can be encrypted and shared with an external server or other participants. During this process, the system (400) can enhance its analysis model by integrating model parameters received from other participants, and the actual data can remain under the control of the client throughout the entire collaboration process. The model integration method and parameter exchange protocol of federated learning are merely examples and are not limited thereto, and may vary depending on the analysis goal and the configuration of the participating clients.

[0186] According to one embodiment, the system (400) may include an AI-based abnormal behavior detection module that analyzes communication patterns with an external server. This module can build a normal behavior model by learning the frequency of normal API calls, the size of data requests, and the form of analysis results. If requests from the external server surge above a certain frequency within a specified time, or if abnormal patterns suspected of being attempts to re-identify a specific individual (e.g., repeatedly requesting analysis under slightly different conditions) are observed, the system (400) can automatically detect this and immediately block the connection or restrict the analysis requests. Additionally, it can proactively respond to potential security threats by notifying the user of relevant warnings. The AI-based abnormal behavior detection model and response policy are merely examples and are not limited thereto, and may vary depending on the system's security policy and operating environment.

[0188] According to one embodiment, the system (400) can utilize a machine learning-based risk prediction model to comprehensively analyze data combinations and external information sources, and can automatically apply additional masking or generalization when the re-identification probability exceeds a threshold, and can optimize the trade-off between analysis accuracy requirements and privacy protection levels in real time. The system (400) can detect indirect identification risks in advance through correlation analysis with external data brokers or public databases, and can predict and preemptively respond to the increase in re-identification risk due to the accumulation of information over time.

[0189] According to one embodiment, the system (400) can record all transformation processes from raw data to the final analysis result on an immutable distributed ledger through a blockchain-based data lineage tracking system. The system (400) can automatically execute data processing rules and privacy policies by utilizing smart contracts, and can cryptographically guarantee integrity by storing hash values ​​and metadata on the blockchain at each data transformation stage. The system (400) is configured so that multiple independent validators can verify the legitimacy of the data processing process through a distributed consensus mechanism, and can also securely store lineage information of large-scale genomic data by linking with a distributed storage system such as IPFS. The blockchain-based lineage tracking method and smart contract implementation technique are merely examples and are not limited thereto, and may vary depending on the blockchain platform and consensus algorithm.

[0190] According to one embodiment, the system (400) can automatically detect and remove hidden personal information or new types of identifiers that are difficult to discover through existing pattern matching using an AI-based intelligent personal information detection engine. The system (400) can extract indirect personal information from genomic annotation information or metadata by utilizing natural language processing and deep learning technologies, and can learn and share new identification patterns discovered in various client environments through federated learning.

[0191] According to one embodiment, the system (400) can completely eliminate the risk of data exposure by performing statistical analysis and machine learning directly in an encrypted state without decrypting non-identified data through homomorphic encryption-based privacy-preserving operations. The system (400) can selectively apply full homomorphic encryption or partial homomorphic encryption depending on the characteristics of the data, and can also execute complex genomic analysis algorithms in the ciphertext domain through encrypted matrix operations and polynomial evaluation. The system (400) can minimize the computational overhead of homomorphic encryption by utilizing circuit optimization and packing techniques, and can use analysis services while completely protecting sensitive genomic data even in a cloud computing environment. The homomorphic encryption implementation method and performance optimization technique are merely examples and are not limited thereto, and may vary depending on the encryption scheme and computational complexity.

[0192] According to one embodiment, the system (400) implements enhanced identity verification by combining fingerprint, iris, voice, facial recognition, etc., through multi-biometric authentication-based data access control, and can prevent data theft or identity impersonation by cross-verifying the consistency between the genomic data itself and the biometric information. The system (400) can prevent system infringement caused by the leakage of a single biometric information by distributing and storing biometric templates and applying a multi-signature technique, and can block attacks using duplicated biometric information through biometric activity detection and anti-spoofing techniques. The multi-biometric authentication configuration method and cross-verification technique are merely examples and are not limited thereto, and may vary depending on biometric recognition technology and security requirements.

[0194] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0195] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0196] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave in order to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0197] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based on the above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

Claims

Claim 1 A genomic analysis system that performs analysis without external leakage of genomic data comprises: a memory for storing instructions; and a processor. When the instructions are executed by the processor, the system stores raw genomic data in client-side storage and applies Differential Privacy technology to inject mathematically calculated noise into the results of data queries so that it is impossible to determine whether a specific individual is included in the dataset, thereby generating non-identifiable data from which personal information has been removed from the raw genomic data. The non-identifiable data is generated in an encrypted form including chromosome numbers, positions, and genetic information, and is configured to have the characteristics of biologically fragmentary and context-removed garbage data, making it impossible to identify individuals or extract meaningful information even if stored. The system utilizes homomorphic encryption technology to encrypt the non-identifiable data, processing it so that an external server can perform analysis without decrypting the ciphertext and return the encrypted result value to the client. After processing, the system transmits the non-identifiable data to an external server to request genetic analysis, and receives an analysis result corresponding to the encrypted result value from the external server. A system that receives and decrypts the encrypted result value only within the client environment to generate a final analysis result in the client environment based on the received analysis result, wherein the system installs a distributed processing module for genomic analysis on a distributed server built within the client environment, and performs the generation of non-identifiable data and processing of analysis results on the distributed server to fundamentally block the path for raw genomic data and source data containing personal information to be exposed outside the client environment, and controls that raw genomic data and personal information are processed only in an isolated internal processing area so as not to access the network transport layer during all communication processes with an external server. Claim 2 In claim 1, the instructions, when executed by the processor, generate non-identifiable data by using a pattern matching algorithm to identify personal information including age, sex, gestational age, BMI (Body Mass Index), and phenotypic information including eye color, skin color, and hair loss status from raw genomic data through a distributed processing module installed in the client environment, and removing it through masking processing; process the non-identifiable data to include information necessary for genetic analysis while preventing external identification of an individual; assign an arbitrary unique number generated in the client environment to the non-identifiable data to manage it so that analysis results can be matched with raw data later; encrypt the non-identifiable data using an encryption algorithm; transmit it to an external server by forming an encrypted secure channel based on HTTPS (Hypertext Transfer Protocol Secure) with TLS (Transport Layer Security) applied; communicate with the external server through web-based communication in the form of a RESTful API (Representational State Transfer Application Programming Interface); and apply the same multi-encryption method in reverse order when receiving analysis result data in the form of unidentifiable code from the external server. A system that decrypts, combines the received analysis result code with raw genome data and personal information stored in the client environment within the client environment to generate personalized genome analysis results, and controls the raw genome data and personal information so that they are not transmitted outside the client environment during the entire analysis process. Claim 3 A genome analysis system according to claim 1, wherein the system is composed of a client storage and an external environment, the client storage stores raw genome data and customer data, generates non-identifiable data by removing personal information from the raw genome data through a distributed processing module within the client storage, transmits the generated non-identifiable data to an analysis module located in the external environment to perform genetic analysis, receives result data of completed analysis from the analysis module, generates a final analysis result by combining the received result data and the customer data at a result processing server within the client storage, and controls the raw genome data and the customer data so as not to be transmitted outside the client storage. Claim 4 In claim 1, the above instructions, when executed by the processor, separate and store raw genomic data and personal information including age, gender, gestational age, BMI, and phenotypic information in a distributed server within the client environment; process the raw genomic data through a distributed processing module installed in the distributed server to generate numerical non-identifiable data that cannot identify an individual and can be interpreted only within the internal system; transmit the non-identifiable data converted into numerical values ​​to an external server to request analysis and receive the result; configure the non-identifiable data to be processed in a form from which original genetic information cannot be inferred or restored so that it has no meaning in itself; assign a random unique number generated through a random number generation algorithm in the client environment to each non-identifiable data to maintain a mapping table within the client that can match the analysis result received from the external server with the raw data; and control the raw genomic data and personal information to be processed only in an isolated internal processing area so as not to access the network transmission layer during all communication processes with the external server. Claim 5 In claim 1, the above instructions constitute a multi-security layer comprising a first de-identification process that de-identifies the data itself when the system transmits non-identified data to an external server when executed by the processor, a second encryption process that encrypts the de-identified data using an encryption algorithm, and a third channel encryption process that transmits the encrypted data through a secure channel; when receiving analysis result data returned after the completion of genetic analysis by the external server, the system receives it in the form of an unidentifiable code composed of a combination of alphabets and numbers and decrypts it by applying the same multi-security layer in reverse order; a result generation module installed in the client environment combines the received code with raw genome data and personal information stored within the client to generate a personalized genome analysis result report, and controls the entire analysis process to proceed within the client environment while the external server performs only analysis operations on the non-identified data. Claim 6 In claim 1, the above instructions, when executed by the processor, utilize the analysis service of an external server via an API call method without installing an analysis engine on the client to respond in real time to version updates of a gene analysis algorithm provided by the external server, access the analysis service through an API authentication token issued according to a service contract, control to block access by invalidating the token upon termination of the contract, and control to process mutation analysis or disease prediction operations on non-identifiable gene sequence data from which personal information has been removed through the API call by utilizing the computing resources of the external server. Claim 7 A system according to claim 1, wherein the instructions, when executed by the processor, perform a procedure to obtain explicit approval for a genome analysis summary result generated in a client environment through a customer's electronic signature or clicking an approval button via a user interface, and display only the approved summary result in a state where it can be extracted from a database and transmitted to an external server, and control the client to modify matters including the selection of data items to be shared, the setting of a sharing period, and restrictions on the purpose of use through a data sharing setting interface, and store the summary result received from the external server as an anonymized dataset to limit its scope of use to machine learning training data for improving the accuracy of an analysis algorithm or for statistical research purposes, and wherein the summary result includes the frequency distribution of individual genetic variations, statistical average values ​​of disease risk, and categorized classification results of drug responsiveness, and unique identifiers capable of identifying an individual are removed. Claim 8 A system according to claim 1 that manages the non-identifiable data by assigning a random unique number generated using a random number generator in a client environment to the non-identifiable data through a mapping table so that the analysis results and raw data can be linked later, encrypts the non-identifiable data using an encryption algorithm including AES (Advanced Encryption Standard) or RSA (Rivest-Shamir-Adleman) as a first encryption, forms a secure channel with TLS (Transport Layer Security) encryption applied to transmit it to an external server, communicates with the external server in JSON or XML data format through web-based communication in the form of a RESTful API, decrypts the data in the form of an unidentifiable code composed of a combination of alphabets and numbers when receiving analysis result data from the external server by applying the same multiple encryption method in reverse order, generates a personalized genomic analysis result by combining the received analysis result code with the corresponding raw genomic data and personal information by referring to the mapping table of the client environment, and controls the transmission of raw genomic data and personal information through a network interface during the entire analysis process so that they are not transmitted outside the client environment. Claim 9 In claim 1, the above instructions, when executed by the processor, are a system that prevents backtracking of original data by injecting noise into variant information of raw genome data during the process of generating non-identifying data, automatically generalizes or removes the relevant part when a rare pattern capable of identifying an individual is detected in the gene sequence to block the risk of re-identification, tracks repetitive analysis requests for the same data and restricts additional analysis if the number of times exceeds a set limit, and controls data usage to enable tracking by recording the generation and transmission history of all non-identifying data transmitted externally in an immutable manner. Claim 10 In claim 1, the above instructions, when executed by the processor, enable the system to perform collaborative analysis while multiple client environments possess their own raw data, wherein an external server distributes an analysis model to each client, each client trains the model with its own non-identifiable data and encrypts only the training results to transmit them to the external server, the external server integrates the received encrypted training results to improve the overall analysis model, ensures that the raw data or non-identifiable data of each client is not transmitted outside the corresponding client environment, evaluates the data quality of each client to exclude data with a reliability level below a specified level from the overall analysis, verifies the accuracy of intermediate results generated during the collaborative analysis process, and controls the system so as not to disclose the actual data content.