Method, apparatus, and system for implementing data visualization tool

The computing device with AI models and GUI enhances data quality assessment by enabling precise visualization and interaction, addressing the limitations of existing methods in analyzing high-dimensional data sets.

WO2026038902A1PCT designated stage Publication Date: 2026-02-19PEBBLOUS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/012348
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-17
Filing Date
2025-08-14
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing data quality assessment methods struggle with precise analysis of unstructured data, fail to reflect high-dimensional data structure, and lack intuitive user interaction for data exploration and quality improvement.

Method used

A computing device equipped with AI models and a GUI enables data set acquisition, selection of processing levels, determination of data properties, and visualization in two or three dimensions, allowing for intuitive interaction and maintaining data structure through vectorization.

Benefits of technology

Facilitates efficient data exploration and manipulation, improving data quality and enhancing machine learning model performance by providing accurate data visualization and interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025012348_19022026_PF_FP_ABST
    Figure KR2025012348_19022026_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment of the present disclosure, a computing apparatus may be provided, the computing apparatus comprising a memory and at least one processor electronically connected to the memory, wherein the at least one processor implemented to execute at least one instruction stored in the memory is configured to carry out the steps of: providing, via a first viewport, a first data image corresponding to a first data set, the first data image including a plurality of data points respectively corresponding to pieces of data included in the first data set; generating, in response to a user input received from a first GUI, first snapshot information including a first scene of the first data image being provided via the first viewport, and providing same via a second viewport; and generating, on the basis of a user input for a second GUI provided via the second viewport, link information for connecting the first snapshot information to an external communication network.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, devices, and systems for implementing data visualization tools

[0001] The present disclosure relates to data processing technology for diagnosing and visualizing data. More specifically, it relates to technology for diagnosing data through data imaging and providing interactive functions by visualizing the data.

[0002] With the widespread adoption of artificial intelligence (AI) and machine learning (ML) technologies across various industries, the importance of technologies capable of effectively processing and analyzing large-scale data sets is increasing. Data quality is a critical factor that directly impacts the performance of AI models (or neural networks), and vectorization and visualization technologies are essential for diagnosing and analyzing data quality. In particular, as large volumes of image, text, and structured data are used to train deep learning-based models, accurately understanding the inherent characteristics of the data is crucial.

[0003] Existing data quality assessment methods primarily rely on integrity verification and basic statistical analysis of structured data, making precise analysis of unstructured data difficult. Furthermore, existing techniques for identifying the inherent distribution of large-scale data sets fail to effectively reflect data structure in high-dimensional spaces and lack the ability to provide visual insights for assessing data quality.

[0004] Existing data visualization technologies also have limitations. Even when using dimensionality reduction techniques to represent high-dimensional data in two- or three-dimensional space, it's difficult to maintain the characteristics of the original data. Furthermore, users lack intuitive interaction when analyzing and manipulating the quality of datasets, making data exploration and quality improvement processes inefficient.

[0005] One task of the present disclosure is to diagnose the quality of a large data set.

[0006] In addition, one task of the present disclosure is to visualize a data set so that it can be effectively expressed while maintaining the inherent structure of the data through a vectorization method for precisely identifying the inherent distribution of a large data set.

[0007] Additionally, one task of the present disclosure is to achieve intuitive user interaction by minimizing the gap between high-dimensional vector space and visualization space.

[0008] According to one embodiment of the present disclosure, a computing device may be provided, comprising a memory and at least one processor electronically connected to the memory, wherein the at least one processor is configured to perform at least one instruction stored in the memory, the computing device being configured to perform the following operations: acquiring a first data set; selecting a first level from among a plurality of levels classified according to a data processing method based on a user input; determining at least one property of a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model; providing, through a first GUI (Graphical User Interface), prior information related to a task of processing the first data set using the first data processing model; and constructing the first data processing model based on a user input to the first GUI.

[0009] In addition, according to one embodiment of the present disclosure, a data processing method may be provided that is set to perform, by at least one processor executing at least one instruction stored in a memory, an operation of acquiring a first data set, an operation of selecting a first level from among a plurality of levels classified according to a data processing method based on a user input, an operation of determining at least one property of a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model, an operation of providing, through a first GUI (Graphical User Interface), prior information related to a task of processing the first data set using the first data processing model, and an operation of constructing the first data processing model based on a user input to the first GUI.

[0010] In addition, according to one embodiment of the present disclosure, an electronic device may be provided, including a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein the at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of receiving a user input for selecting a first level among a plurality of levels classified according to a data processing method, an operation of displaying a first GUI (Graphic User Interface) through the display that indicates prior information related to a task of processing the first data set using a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model, an operation of obtaining a diagnosis result for the first data set based on a first vector set defined in a first-dimensional embedding area from the first data processing model, and an operation of providing the first vector set to a first visualization tool to display a two-dimensional or three-dimensional first data image, the first data image corresponding to the first data set, through the display.

[0011] In addition, according to one embodiment of the present disclosure, by at least one processor executing at least one instruction stored in a memory, an operation of obtaining a data set, an operation of embedding the data set in a specific dimension to obtain a vector set, the vector set including a plurality of vectors corresponding to a plurality of data included in the data set, an operation of screening the vector set by calculating at least one feature value based on the vectors included in the vector set using a screener having at least one metric set, an operation of identifying at least one vector whose at least one feature value satisfies a predetermined condition, an operation of obtaining a data image including a plurality of data points representing the data set in two dimensions or three dimensions by processing the vector set using at least one visualization tool, an operation of determining a first area including at least one data point corresponding to the at least one vector on the data image and generating tag information associated with the first area, and a first scene including the first area, the first scene being generated by capturing the first area from a specific view-point, and including the tag information. A method may be provided that includes an action to store snapshot information.

[0012] In addition, according to one embodiment of the present disclosure, by at least one processor executing at least one instruction stored in the computer program, the computer program includes an operation of acquiring a data set, an operation of embedding the data set in a specific dimension to acquire a vector set, the vector set including a plurality of vectors corresponding to a plurality of data included in the data set, an operation of screening the vector set by calculating at least one feature value based on the vectors included in the vector set using a screener having at least one metric set, an operation of identifying at least one vector whose at least one feature value satisfies a predetermined condition, an operation of processing the vector set using at least one visualization tool to acquire a data image including a plurality of data points representing the data set in two dimensions or three dimensions, an operation of determining a first region including at least one data point corresponding to the at least one vector on the data image and generating tag information associated with the first region, and a first scene including the first region, the first scene being generated by capturing the first region from a specific view-point, and the tag information. A computer program may be provided that is configured to perform an operation of storing snapshot information including:

[0013] In addition, according to one embodiment of the present disclosure, a computing device may be provided, comprising a memory and at least one processor electronically connected to the memory, wherein the at least one processor is configured to execute at least one instruction stored in the memory, and is configured to perform an operation of providing a first data image corresponding to a first data set through a first view port, the first data image including a plurality of data points corresponding to each of data included in the first data set; an operation of generating first snapshot information including a first scene for the first data image being provided through the first view port in response to a user input received through a first GUI and providing the first snapshot information through a second view port; and an operation of generating link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.

[0014] In addition, according to one embodiment of the present disclosure, a data interaction method may be provided, including an operation of providing a first data image corresponding to a first data set, the first data image including a plurality of data points corresponding to each of data included in the first data set, through a first view port by at least one processor executing at least one instruction in a memory, an operation of generating first snapshot information including a first scene for the first data image being provided through the first view port in response to a user input received through a first GUI and providing the first snapshot information through a second view port, and an operation of generating link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.

[0015] In addition, according to one embodiment of the present disclosure, an electronic device may be provided, comprising a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of providing a first data image corresponding to a first data set, the first data image including a plurality of data points corresponding to each of data included in the first data set, through a first view port of the display, an operation of providing first snapshot information including a first scene for the first data image being provided through the first view port through the display in response to a user input received through the first GUI, and an operation of providing link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.

[0016] In addition, according to one embodiment of the present disclosure, a method may be provided, including an operation of visualizing a data set by at least one processor executing at least one instruction stored in a memory to provide a first data image corresponding to the data set, an operation of receiving a user input for a first area on the first data image, an operation of determining a latent code corresponding to the first area and providing the determined latent code to a generation model to generate synthetic data, an operation of inputting the synthetic data to a data processing model, the data processing model being trained to embed data into a specific dimension, to obtain a synthetic vector corresponding to the synthetic data, and an operation of visualizing the synthetic vector to provide a synthetic point corresponding to the synthetic data on the first data image.

[0017] In addition, according to one embodiment of the present disclosure, a computing device may be provided, including a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein the at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of visualizing a data set and providing a first data image corresponding to the data set through the display, an operation of receiving a user input for a first area on the first data image, an operation of determining a latent code corresponding to the first area and providing the determined latent code to a generation model to generate synthetic data, an operation of inputting the synthetic data into a data processing model, the data processing model being trained to embed data into a specific dimension, to obtain a synthetic vector corresponding to the synthetic data, and an operation of visualizing the synthetic vector and providing a synthetic point corresponding to the synthetic data on the first data image through the display.

[0018] The solutions to the problems of the present invention are not limited to the solutions described above, and solutions not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from this specification and the attached drawings.

[0019] According to one embodiment of the present disclosure, a computing device can automatically diagnose the quality of a large data set and intuitively visualize it. This allows data analysts, AI model developers, and researchers to more easily understand the inherent structure of the data and develop strategies to improve data quality.

[0020] Additionally, according to one embodiment of the present disclosure, a user can more efficiently explore and manipulate data during the data visualization process, and utilize it to improve the performance of a machine learning model.

[0021] The effects of the present disclosure are not limited to the effects described above, and effects not mentioned will be clearly understood by those skilled in the art to which the present invention pertains from this specification and the attached drawings.

[0022] FIG. 1 is a diagram illustrating a configuration of a computing device according to various embodiments.

[0023] FIG. 2 is a diagram illustrating various data processing methods included in a data clinic service provided by a computing device according to various embodiments.

[0024] FIG. 3 is a diagram illustrating various systems for providing data clinic services according to various embodiments, and artificial intelligence models and algorithms for constructing the systems.

[0025] FIG. 4 is a diagram illustrating a method for a computing device to provide a data image according to various embodiments.

[0026] FIG. 5 is a diagram illustrating a method for a computing device to obtain characteristics of a data set according to various embodiments.

[0027] FIG. 6 is a diagram illustrating a data lens processing system and a data imaging system according to various embodiments.

[0028] FIG. 7 is a diagram illustrating a system in which a computing device visualizes data and provides user interaction functions, according to various embodiments.

[0029] FIG. 8 is a diagram illustrating a data visualization and interaction method according to various embodiments.

[0030] FIG. 9 is a drawing for explaining detailed steps performed in an imaging application step according to various embodiments.

[0031] FIG. 10 is a diagram illustrating the results of data imaging and diagnosis according to the level of the lens and the visualization tool according to various embodiments.

[0032] FIG. 11 is a diagram illustrating a lens builder included in a computing device according to various embodiments.

[0033] FIG. 12 is a diagram illustrating a method for a computing device to build a data processing model for data imaging and diagnosis based on user input, according to various embodiments.

[0034] FIG. 13 is a diagram for explaining information included in dictionary information according to various embodiments.

[0035] FIG. 14 is a diagram illustrating an example of a computing device constructing a first data processing model according to various embodiments.

[0036] FIG. 15 is a diagram illustrating another example of a computing device constructing a first data processing model according to various embodiments.

[0037] FIG. 16 is a diagram illustrating another example of a computing device constructing a first data processing model according to various embodiments.

[0038] FIG. 17 is a diagram illustrating a method of imaging and visualizing a data set using a first data processing model built on a computing device according to various embodiments.

[0039] FIG. 18 is a diagram illustrating a method for a computing device to recommend a visualization tool according to various embodiments.

[0040] FIG. 19 is a diagram illustrating a method for a computing device to generate synthetic data based on data imaging according to various embodiments.

[0041] FIG. 20 is a diagram illustrating a method for a computing device to generate synthetic data based on user input according to various embodiments.

[0042] FIG. 21 is a diagram illustrating an example of a method for a computing device to generate synthetic data based on user input, according to various embodiments.

[0043] FIG. 22 is a diagram illustrating a method for a computing device to remove data based on user input, according to various embodiments.

[0044] FIG. 23 is a diagram illustrating a method for a computing device to perform data improvement in response to a data improvement request and provide visual interaction therefor, according to various embodiments.

[0045] FIG. 24 is a diagram illustrating an example of a computing device in which a data processing method including a snapshot function is implemented, according to various embodiments.

[0046] FIG. 25 is a diagram illustrating a method for a computing device to screen a data set to generate snapshot information, according to various embodiments.

[0047] FIG. 26 is a diagram illustrating a method for a computing device to provide snapshot information based on user input according to various embodiments.

[0048] Figure 27 is an example of a screen provided by a computing device.

[0049] FIG. 28 is a diagram illustrating a function of a computing device to reproduce snapshot information according to various embodiments.

[0050] FIG. 29 is a diagram illustrating a method for a computing device to generate diagnostic reports and snapshot information in conjunction with each other, according to various embodiments.

[0051] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In describing the embodiments, descriptions of technical details that are well-known in the technical field to which the present disclosure pertains and are not directly related to the present disclosure will be omitted. This is to avoid obscuring the gist of the present disclosure by omitting unnecessary explanations and to convey the gist more clearly.

[0052] Since the embodiments described in this specification are intended to clearly explain the idea of ​​the present invention to a person having ordinary skill in the art to which the present invention pertains, the present invention is not limited to the embodiments described in this specification, and the scope of the present invention should be interpreted to include modified or altered examples that do not depart from the idea of ​​the present invention.

[0053] The terms used in this specification have been selected from widely used terms, taking into account the functions of the present invention. However, these terms may vary depending on the intentions of those skilled in the art, precedents, or the emergence of new technologies. However, if a specific term is defined and used with an arbitrary meaning, the meaning of that term will be described separately. Therefore, the terms used in this specification should be interpreted based on the actual meaning of the term and the overall content of this specification, rather than simply the name of the term.

[0054] The drawings attached to this specification are intended to facilitate explanation of the present invention, and the shapes depicted in the drawings may be exaggerated as necessary to help understanding of the present invention, and therefore the present invention is not limited by the drawings.

[0055] In this specification, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in that phrase, or all possible combinations thereof.

[0056] If a detailed description of the composition or function of a known disclosure related to the present invention in this specification is deemed to obscure the gist of the present invention, a detailed description thereof will be omitted as necessary. Furthermore, the numbers (e.g., "first," "second," etc.) used throughout the description of this specification are merely identifiers used to distinguish one component from another.

[0057] In addition, the suffixes "part" and "part" for components used in the following description are given or used interchangeably only for the convenience of writing the specification, and do not have distinct meanings or roles in themselves.

[0058] That is, the embodiments of the present disclosure are provided to make the present disclosure complete and to inform those skilled in the art of the scope of the present disclosure, and the invention of the present disclosure is defined solely by the scope of the claims. Like reference numerals refer to like elements throughout the specification.

[0059] Terms such as “first” and / or “second” may be used to describe various components, but the components should not be limited by the terms. The terms are only for the purpose of distinguishing one component from another, for example, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component, without departing from the scope of the present disclosure.

[0060] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between. Conversely, when a component is referred to as being "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Other expressions that describe the relationship between components, such as "between" and "directly between" or "adjacent to" and "directly adjacent to", should be interpreted similarly.

[0061] Each block of the flowchart drawings and combinations of flowchart drawings in the drawings can be performed by computer program instructions. These computer program instructions can be installed in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, so that the instructions executed by the processor of the computer or other programmable data processing equipment create a means for performing the functions described in the flowchart block(s). These computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to implement the functions in a specific manner, so that the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes an instruction means for performing the functions described in the flowchart block(s). Since the computer program instructions may be installed on a computer or other programmable data processing device, a series of operational steps may be performed on the computer or other programmable data processing device to create a computer-executable process, so that the instructions that cause the computer or other programmable data processing device to perform the steps for performing the functions described in the flowchart block(s) may also be able to provide steps for performing the functions described in the flowchart block(s).

[0062] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0063] Additionally, each block may represent a module, segment, or portion of code that contains one or more executable instructions for performing a specified logical function(s). It should also be noted that in some alternative implementation examples, the functions mentioned in the blocks may occur out of order. For example, two blocks shown in succession may in fact be performed substantially concurrently, or the blocks may sometimes be performed in reverse order, depending on their respective functions. For example, the operations performed by a module, program, or other component may be performed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be performed in a different order, omitted, or one or more additional operations may be added.

[0064] The term 'unit' as used in this disclosure means a software or hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC). The 'unit' performs specific roles, but is not limited to software or hardware. The 'unit' may be configured to reside on an addressable storage medium and may be configured to play one or more processors. Accordingly, according to some embodiments, the 'unit' includes components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided within the components and 'units' may be combined into a smaller number of components and 'units' or further separated into additional components and 'units'. Additionally, the components and '~parts' may be implemented to activate one or more CPUs within a device or secure multimedia card. Furthermore, according to various embodiments of the present disclosure, the '~parts' may include one or more processors.

[0065] The operating principles of the present disclosure are described in detail below with reference to the attached drawings. In the following description of the present disclosure, detailed descriptions of related known functions or configurations will be omitted if they are deemed to unnecessarily obscure the gist of the present disclosure. Furthermore, the terms described below are defined based on the functions of the present disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the overall content of this specification.

[0066] FIG. 1 is a diagram illustrating a configuration of a computing device according to various embodiments.

[0067] Referring to FIG. 1, a computing device (e.g., an electronic device including a computing means such as a server or client device, hereinafter referred to as a “computing device”) (100) according to one embodiment may include a processor (110), a memory (120), a storage device (130), a communication circuit (140), and a bus (not shown). The configuration of the computing device (100) is not limited to the configuration illustrated in FIG. 1 or the configuration described above, and may further include hardware or software configurations included in general computing devices or mobile devices.

[0068] The processor (110) may include at least one processor, at least some of which are implemented to provide different functions. For example, the processor (110) may execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of the computing device (100) connected to the processor (110) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (110) may store instructions or data received from other components in the memory (120) (e.g., a volatile memory), process the instructions or data stored in the volatile memory, and store the resulting data in the non-volatile memory. According to one embodiment, the processor (110) may include a main processor (e.g., a central processing unit or an application processor) or an auxiliary processor (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, if the computing device (100) includes a main processor and a secondary processor, the secondary processor may be configured to use less power than the main processor or to be specialized for a given function. The secondary processor may be implemented separately from the main processor or as a part thereof. The secondary processor may control at least a portion of functions or states associated with at least one component of the computing device (100) (e.g., a display (240) or a communication circuit), for example, on behalf of the main processor while the main processor is in an inactive (e.g., sleep) state, or together with the main processor while the main processor is in an active (e.g., application execution) state.In one embodiment, the auxiliary processor (e.g., an image signal processor or a communication processor) may be implemented as part of another functionally related component (e.g., a communication circuit). In one embodiment, the auxiliary processor (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. Meanwhile, the operation of the computing device (100) described below may be understood as the operation of the processor (110).

[0069] According to various embodiments, the memory (120) may include at least one memory, at least some of which are implemented to provide different functions. The memory (120) may store various data used by at least one component (e.g., the processor (110)) of the computing device (100). The data may include, for example, software (e.g., a program) and input data or output data for instructions related thereto. The memory (120) may include volatile memory or non-volatile memory. The memory (120) may be implemented to store an operating system, middleware or applications, and / or the artificial intelligence model described above.

[0070] Additionally, the memory (120) may include a plurality of instructions (121) that direct the operations of the processor (110) to implement the functions provided by the service. At this time, the processor (110) may execute at least some of the plurality of instructions stored in the memory (120). The computing device (110) may include a software server including the processor (110) that executes the functions provided by the service based on at least some of the plurality of instructions.

[0071] The storage device (130) can provide a mass storage device to the computing device (100). The storage device (130) can be a computer-readable medium. For example, the storage device (130) can be a floppy disk device, a hard disk device, an optical disk device, a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other configuration device. In addition, a computer program product is explicitly embodied in an information medium. The computer program product includes instructions that, when executed, perform one or more methods as described above. The information medium is a computer-readable medium or a machine-readable medium, such as the memory (120), the storage device (130), or the memory of the processor (110).

[0072] Additionally, the storage device (130) may include a database (DB). The storage device (130) may include a database having a pre-structured data structure. The computing device (110) may store data sets having interrelated relationships in the database.

[0073] A computing device according to the present disclosure can provide services based on various artificial intelligence frameworks performed by at least one processor and a memory electronically connected to at least one processor.

[0074] In this regard, the memory (120) or storage device (130) may store at least one artificial intelligence model implementing various types of artificial intelligence (or machine learning) frameworks that can be trained to perform a given task. For example, support vector machines, decision trees, neural networks, etc. are just a few examples of machine learning frameworks used in various applications such as image processing and natural language processing. Some artificial intelligence frameworks, such as neural networks, may utilize layers of nodes that perform specific operations.

[0075] In a neural network, nodes are connected to each other through one or more edges. A neural network may include an input layer, an output layer, and one or more intermediate layers. Each node may process its inputs according to a predefined function and provide output to subsequent layers, or in some cases, previous layers. The input to a particular node may be multiplied by a weight value corresponding to the edge between the input and the node. Additionally, each node may have a separate bias value used to generate the output. Various learning procedures can be applied to learn the edge weights and / or bias values ​​(parameters).

[0076] A neural network architecture may have multiple layers that perform different specific functions. For example, one or more layers of nodes may collectively perform specific operations, such as pooling, encoding, or convolution operations. As used herein, the term "layer" may refer to a group of nodes that share inputs and outputs, such as communicating with external sources or other layers of the network. The term "calculation" may refer to a function that can be performed by one or more layers of nodes. The term "model structure" may refer to the overall architecture of a layered model, including the number of layers, the connectivity of the layers, and the types of operations performed by individual layers. The term "neural network structure" may refer to the model structure of a neural network. The terms "trained model" and / or "tuned model" may refer to the model structure along with the parameters for the trained or tuned model structure. For example, two trained models may have different values ​​for parameters even though they share the same model structure, such as when they are trained on different training data or when the training process has an underlying probabilistic process.

[0077] "Transfer learning" is a broad approach for training models with limited task-specific training data for a specific task. In transfer learning, a model is first pretrained on another task for which valuable training data is available, and then adapted to a specific task using task-specific training data.

[0078] The term "pre-training," as used herein, refers to training a model on a pre-training dataset to adjust model parameters in a manner that allows subsequent adjustments of those model parameters to tailor the model to one or more specific tasks. In some cases, pre-training may involve a self-supervised learning process on unlabeled training data, where the "self-supervised" learning process involves learning from the structure of pre-training examples in the absence of explicit (e.g., manually provided) labels. Subsequent modification of the model parameters obtained through pre-training is referred to herein as "tuning." Tuning may be performed for one or more tasks using supervised learning on explicitly labeled training data, and in some cases, a task different from pre-training may be used for tuning.

[0079] A communication bus (not shown) may be a configuration for electronically (or communicatively) connecting multiple components included in a computing device. That is, each component may be interconnected using various buses and mounted on a common motherboard or in another suitable manner.

[0080] The input / output interface (not shown) may include an input interface that is connected to an input device and receives an input signal, or an output interface that is connected to an output device and outputs an output signal.

[0081] Additionally, the computing device (100) may further include at least one communication circuit (140) for communicating with an external device.

[0082] The communication circuit (140) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the computing device (100) and an external computing device, and the performance of communication through the established communication channel. The communication circuit may operate independently from the processor (110) (e.g., a program processor) and may include one or more communication processors (e.g., communication chips) that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication circuit (140) may include a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external computing device via a first network (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module can identify or authenticate the computing device (100) within a communication network such as the first network or the second network by using subscriber information stored in the subscriber identification module (e.g., an international mobile subscriber identity (IMSI)). The wireless communication module can support a 5G network subsequent to a 4G network and next-generation communication technologies, such as new radio access technology (NR).NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high-reliability and low-latency communications (URLLC (ultra-reliable and low-latency communications)). The wireless communication module can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module can support various technologies to secure performance in the high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module can support various requirements specified in a computing device (100), an endoscope device, or a network system. According to one embodiment, the wireless communication module can support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL) each, or 1 ms or less for round trip) for URLLC implementation.

[0083] The computing device (100) may be implemented to include at least some of the above-described components (processor, communication circuitry, memory, display). For example, a user device may be implemented to include a processor, communication circuitry, memory, sensors, and a display. Additionally, for example, a server device may be implemented to include a processor, communication circuitry, and memory.

[0084]

[0085] FIG. 2 is a diagram illustrating various data processing methods included in a data clinic service provided by a computing device according to various embodiments.

[0086] Referring to FIG. 2, the data clinic service may include various data processing methods. These various data processing methods may be encoded and stored in the memory of the computing device, and at least one processor included in the computing device may be configured to execute at least one encoded instruction. Specifically, at least one processor may process a received input data set based on the various data processing methods and output an output data set.

[0087] For example, a computing device according to various embodiments of the present disclosure may perform, but is not limited to, an operating method for data imaging, an operating method for data enhancement, an operating method for data generation, an operating method for data feature extraction, or an operating method for data evaluation.

[0088] Additionally, each of the above-described operating methods can be performed based on operating algorithms of at least one processor included in the computing device.

[0089] For example, a computing device according to various embodiments of the present disclosure may perform, but is not limited to, a data imaging algorithm, a data enhancement algorithm, a data generation algorithm, a data feature extraction algorithm, or a data evaluation algorithm.

[0090] At this time, the names of each operation method and algorithm are arbitrarily named according to the output results for the convenience of explanation, so each operation method or algorithm is only defined based on the operations performed by the processor, and the name of the operation method or algorithm itself does not limit the invention.

[0091] More specifically, a computing device according to various embodiments of the present disclosure can process an input data set according to a data imaging algorithm to generate an image for the input data set.

[0092] Additionally, a computing device according to various embodiments of the present disclosure can improve data by processing an input data set according to a data improvement algorithm, and can generate a result of the improvement.

[0093] Additionally, a computing device according to various embodiments of the present disclosure can generate synthetic data by processing an input data set according to a data generation algorithm.

[0094] Additionally, a computing device according to various embodiments of the present disclosure can process an input data set according to a data feature extraction algorithm to extract a property of the input data set.

[0095] Additionally, a computing device according to various embodiments of the present disclosure can process an input data set according to a data evaluation algorithm to evaluate the quality of the input data set.

[0096] Details of each of the above algorithms are explained below.

[0097] Furthermore, the computing device according to various embodiments of the present disclosure can perform the various operation methods or algorithms described above in parallel, sequentially, or selectively. Specifically, the computing device can use the same input data as input values ​​for different algorithms in parallel, can use the result values ​​output by a specific algorithm sequentially as input values ​​for another algorithm, or can selectively perform some of the algorithms among a plurality of algorithms according to a predetermined method.

[0098] In addition, the various operation methods or algorithms for the data clinic described above can be performed in a deep learning model included in a computing device according to various embodiments of the present disclosure. Specifically, the computing device according to various embodiments of the present disclosure may include one deep learning model for performing the various operation methods or algorithms described above, but is not limited thereto, and may include multiple deep learning models for performing each of the operation methods or algorithms described above, or may include one or more deep learning models for performing at least some of the various operation methods or algorithms described above.

[0099] Figure 3 is a diagram illustrating various systems for providing data clinic services according to various embodiments, as well as artificial intelligence models and algorithms for constructing such systems. Here, a system may refer to a system that includes at least one software or hardware configuration to perform a specific function.

[0100] Referring to FIG. 3, a computing device according to the present disclosure may include a data clinic system composed of various artificial intelligence (or neural network, machine learning, etc.) models to provide clinic services.

[0101] For example, the computing device may include, but is not limited to, a data imaging system, a data diagnostic system, and a data treatment system.

[0102] Here, the data imaging system may include, but is not limited to, a lens processing model for determining an optimal dimension for representing the characteristics of the data, an imaging model for obtaining a data image reflecting the inherent characteristics of the data, or a visualization model for visually representing the data.

[0103] Additionally, the data diagnosis system may include, but is not limited to, a diagnosis model for diagnosing at least one characteristic of data or a quality assessment model for evaluating the quality of data.

[0104] Additionally, the data treatment system may include, but is not limited to, a synthetic model (or generation model) for generating targeted virtual data (or synthetic data) as needed, a data diet model for removing at least a portion of the data, or a data correction model for adjusting the characteristics of at least a portion of the data.

[0105] Various machine learning models included in a computing device may be comprised of multiple modules stored in memory. In the present disclosure, a module may include multiple hardware components for implementing an artificial intelligence model that performs a specific function. For example, a module may include, but is not limited to, an encoder, a decoder, a generator, a discriminator, an adapter, a natural language processing module, or a large language model (LLM).

[0106] The computing device can store the plurality of modules described above, and can construct an AI framework based on at least some of the modules to obtain an AI model for a data clinic. For example, a data lens included in a data imaging system can be implemented as an AI model including at least one encoder or at least one adapter, but is not limited thereto.

[0107] FIG. 4 is a diagram illustrating a method for a computing device to provide a data image according to various embodiments.

[0108] Referring to FIG. 4, a computing device according to various embodiments of the present disclosure may receive a data set and provide an Image of Data (IOD).

[0109] At this time, the data set may be data of dimension M (M>0). In other words, the data set may be a data set defined on an M-dimensional input space (310).

[0110] Additionally, the data set may be a single-modality data set. For example, the data set may be an image data set. Additionally, the data set may be a text data set. Furthermore, without limitation, the data set may be a collection of data with different modalities. For example, the data set may be an image data set including annotation information. Additionally, the data set may be a mixed data set of images and text.

[0111] A computing device according to various embodiments of the present disclosure can receive and process data of all modalities that can be used for deep learning, such as time series data sets, sensor data sets, as well as the image data and text data described above, as input data sets.

[0112] The data image (IOD) provided by the computing device according to various embodiments of the present disclosure may be an image that processes an input data set and displays it in an imaging space (320). Here, the image does not mean a 2D image, but is a general expression that visually represents data. Specifically, the imaging space (320) is a concept that includes a 2D space, a 3D space, and an N-dimensional virtual space, and means a space in which a data image provided according to an embodiment appears. For example, when the computing device processes the input data set and outputs the data image in PDF format, it may output an output that displays the data image in a 2D or 3D imaging space, but is not limited thereto.

[0113] When a computing device according to various embodiments of the present disclosure includes an output device (not shown), the computing device can provide a data image through the output device. For example, the computing device can provide the data image by outputting the data image through a display. In this case, the imaging space (320) may be a screen of a display. Furthermore, for example, the computing device can provide the data image by outputting the data image through a printing device. In this case, the imaging space (320) may be paper output by the printing device.

[0114] Additionally, when a computing device according to various embodiments of the present disclosure communicates with an external device via a communication unit, the computing device may provide a data image via the external device. In this case, the imaging space (320) may be a display screen of the external device. For example, when the computing device is a server device, the server device may provide the data image by transmitting the data image to at least one external device that communicates with the server device via a network connected to the server device.

[0115] A computing device according to various embodiments of the present disclosure can obtain a data image based on a vector set (or data point set, point data set, etc.) (330) corresponding to an input data set.

[0116] At this time, the computing device can obtain a vector set by mapping the data included in the input data set to an embedding space (or latent space) of a specific dimension. Specifically, the computing device can obtain a vector set by identifying a manifold formed by the data set in an embedding space of a specific dimension. Here, the manifold may refer to a shape that the input data set represents in an embedding space of a specific dimension. In other words, the manifold may refer to an area where a vector set is identified or a shape formed by a vector set when mapping the input data set to a vector set in an embedding space of a specific dimension.

[0117] An IOD (Information Object Descriptor) may be a data set that visualizes each data point within the data set. In this case, the shape or color of the visualized point may vary depending on the embodiment, and therefore the term "point" itself is not intended to limit the invention. Furthermore, a point may be expressed using various terms depending on the embodiment. For example, a point may be expressed using terms such as a vector or feature appearing in an embedding space or latent space, but is not limited thereto.

[0118] In order for a computing device according to various embodiments of the present disclosure to provide a data image, as described above, it is necessary to identify a vector set corresponding to the input data set.

[0119] The computing device can obtain a vector set by mapping the data set to an N-dimensional embedding space based on a predefined condition defined by a mapping function (e.g., a pre-stored matrix for mapping to an embedding space of a specific dimension). For example, the computing device can obtain a vector set by encoding the data set, but is not limited thereto. For example, the computing device can input the data set to a pre-trained encoder and obtain a vector set through the output layer of the encoder, but is not limited thereto.

[0120] Here, embedding refers to the process of converting high-dimensional data into low-dimensional vectors, preserving similarities and structural relationships between data points. This embedding is performed in a way that preserves the core information of the data while increasing computational efficiency. During the embedding process, each data point is represented as a vector, and these vectors can reflect the distribution and characteristics of the entire data set. This provides a foundation for analyzing the statistical characteristics and inherent patterns of the data set.

[0121] A computing device may include a data lens (400) for obtaining a data image (IOD) by embedding and visualizing a data set. At this time, the data lens (400) may include at least one processing configuration for processing data. Specifically, the data lens (400) may include at least one neural network model (e.g., an encoder, etc.) for obtaining a vector set based on the data set and at least one visualization model (e.g., PCA, T-SNE, UMAP, etc.) for visualizing the data set based on the vector set to obtain a data image. Specifically, the data lens (400) may obtain a vector set corresponding to the data set by embedding the data set in an N-dimensional latent space, and may obtain a data image (IOD) corresponding to the data set by representing the vector set in an M-dimensional (e.g., 2-dimensional or 3-dimensional) imaging space (320).

[0122]

[0123] FIG. 5 is a diagram illustrating a method for a computing device to obtain characteristics of a data set according to various embodiments.

[0124] Referring to FIG. 5, the computing device can process the acquired data set to obtain characteristic information corresponding to the data set.

[0125] The properties of a data set or data may include information related to the distribution (e.g., geometric distribution or statistical distribution) of the data set or data. Specifically, the properties may include the property values ​​of each data included in the data set. For example, a computing device may obtain property information indicating the distribution of the property values ​​of the data included in the data set. Furthermore, the computing device may obtain the property information of the data set based on the statistical distribution, such as the mean, deviation, or variance, of the property values ​​of each data.

[0126] For example, the properties of a data set or data may include intrinsic characteristics related to the distribution of the data set itself. For example, the properties of a data set or data may include, but are not limited to, the density, homogeneity, bias, or distribution of the data set or data.

[0127] As another example, the properties of a data set or data may include task-dependent properties related to the task for which the data set is utilized (e.g., classification). For example, the properties of a data set or data may include, but are not limited to, the labeling error rate or the proportion of data pairs that are geometrically adjacent (hard-negative) but belong to different classes.

[0128] Additionally, the computing device may store computational metrics corresponding to each characteristic of the data set or data in memory. More specifically, the computing device may store, but is not limited to, metrics for computing the density of the data set or data, metrics for computing the homogeneity of the data set or data, metrics for computing the bias of the data set or data, metrics for computing the distribution of the data set or data, etc.

[0129] Additionally, the computing device can acquire data set characteristics based on stored operational metrics, using a data feature extraction algorithm built using an artificial neural network. Specifically, the feature extraction algorithm can be implemented using a feed-forward neural network.

[0130] For example, the computing device may include, but is not limited to, a separate neural network for computing characteristics of a data set, or may include a neural network including layers for computing characteristics of a data set.

[0131] For example, a computing device may include an artificial neural network for feature extraction designed to extract features of a data set when inputted with the data set. The artificial neural network for feature extraction may be an artificial neural network that has undergone transfer learning to compute the features of the data.

[0132] As another example, a computing device can acquire characteristics of a data set by constructing an artificial neural network that adds a layer for extracting data characteristics to a neural network model (e.g., a data lens, an encoder, etc.) for providing data images based on the data set. Specifically, the computing device can identify a vector set based on the data set and acquire characteristics of the data set or data based on the identified vector set.

[0133] At this time, the computing device can obtain the characteristic value of each data included in the data set by processing each vector included in the vector set with a predetermined algorithm. In this case, the computing device can calculate the characteristic value based on the geometric distribution or statistical distribution of each vector included in the vector set, and can assign the calculated characteristic value to the corresponding data. At this time, the characteristic value can be calculated based on the distance between vectors. For example, the characteristic value can be obtained based on the number of vectors existing within a predetermined distance from a specific vector (or data point), but is not limited thereto. In addition, for example, the characteristic value can be obtained based on the average value of the distances from a specific vector to a predetermined number of nearby vectors, but is not limited thereto. For example, the computing device can calculate the average distance value based on the distance values ​​from a specific vector to K nearby vectors, and can obtain the first characteristic value (e.g., density, etc.) of the specific vector based on the calculated average distance value, but is not limited thereto.

[0134] In order to optimize the framework of an artificial intelligence model and produce highly accurate results, the quality of the data used to train the model is very important.

[0135] As mentioned above, data quality is a concept that includes both quantitative and qualitative quality. Therefore, for successful training of an artificial intelligence model, it is necessary to (i) secure a sufficient amount of training data to train the artificial intelligence model, (ii) secure training data with high-quality inherent characteristics (e.g., unbiased distribution), and (iii) secure training data with characteristics (e.g., task-dependent properties) appropriate for the training purpose (e.g., task of the artificial intelligence model).

[0136] A computing device according to one embodiment of the present disclosure can synthesize, modify (or adjust), or remove data in a way that enhances the inherent characteristics and task-dependent characteristics of the data set to obtain high-quality learning data.

[0137] In addition, the computing device according to the present disclosure can improve the overall quality of a data set by removing at least some data from the data set.

[0138] Typically, data downsampling or undersampling techniques are used to address data imbalance. However, existing undersampling methods have the disadvantage of negatively impacting the learning performance of machine learning models, as they remove data without considering the characteristics of the machine learning dataset.

[0139] A computing device according to the present disclosure can improve the learning efficiency of an artificial intelligence model that is learned by appropriately removing at least some data from a data set.

[0140] FIG. 6 is a diagram illustrating a data lens processing system and a data imaging system according to various embodiments.

[0141] Referring to (a) of FIG. 6, a computing device may acquire a data lens system based on a data set. At this time, the data lens system may be a term meaning at least one configuration for mapping the data set to a specific embedding space. For example, the data lens system may include at least one encoder for mapping the data set to an embedding space (or latent space) of a specific dimension and / or at least one adapter for adjusting parameters. In addition, for example, the data lens system may include a neural network layer composed of at least one node for identifying a latent variable (or latent feature vector) corresponding to the data set.

[0142] The computing device (3500) can determine a data lens system that processes the data set to preserve inherent characteristics of the data set based on the data set.

[0143] For example, a computing device can acquire a lens system corresponding to a data set based on a database. Specifically, the computing device can search the database for a lens system corresponding to the input data set based on the characteristics of the input data set.

[0144] As another example, a computing device can obtain a lens system corresponding to a data set based on a lens processing algorithm. Specifically, the computing device can calculate the optimal dimensionality that preserves the inherent characteristics of the input data set.

[0145] Referring to (b) of FIG. 6, a computing device may process a data set based on a data imaging system (3510) including a determined data lens system to obtain a data image. At this time, the imaging system (3510) may include a lens system composed of at least one module (e.g., an encoder, an adapter, etc.).

[0146] In this case, the computing device can obtain a data image representing the intrinsic characteristics of the data set using the imaging system (3510).

[0147] This disclosure provides a method for precisely identifying the inherent distribution of data by vectorizing it and intuitively visualizing it in two- or three-dimensional space, enabling the effective analysis and visualization of large data sets. Specifically, interaction and UI / UX implementation technologies are applied to minimize the gap between high-dimensional vector space and visualization space, enabling users to more easily understand and manipulate data quality.

[0148] Furthermore, the method of the present disclosure can receive large data sets as input and provide a clearer visual representation while preserving the inherent characteristics of the data set (e.g., geometric distribution or correlations between data). This allows users to intuitively understand the data distribution and can be effectively utilized to improve data quality during AI learning or big data analysis.

[0149] FIG. 7 is a diagram illustrating a system in which a computing device visualizes data and provides user interaction functions, according to various embodiments.

[0150] Referring to FIG. 7, the computing device (700) can image the data set (710) and obtain a vector set corresponding to the data set (710). At this time, the computing device (700) can obtain a data image (720) based on the vector set. The data image (720) may be data that visualizes the data set (710) in two dimensions or three dimensions. The computing device (100) can identify an N-dimensional (N>3)-dimensional vector set based on the data set (710) and obtain the data image (720) by reducing the dimensionality of the N-dimensional vector set to two dimensions or three dimensions. That is, the computing device (700) can express the data image (720) corresponding to the data set (700) in a two-dimensional imaging space or a three-dimensional imaging space.

[0151] The computing device (700) can provide a data image (720) through a network environment (e.g., a web or app environment). The user device (701) can check the data image (720) through the network environment and obtain information about the checked data image (720). At this time, the user device (701) may be an electronic device such as a laptop computer, a smartphone, a tablet, etc. In addition, for example, the user device (701) may be a wearable device such as a smart watch, smart glasses, smart glasses, etc. In addition, for example, the user device (701) may be a media device such as a streaming media device, a media player, an automobile entertainment system, etc. In addition, for example, the user device (701) may be an XR (mixed reality) device such as a device including a VR device or AR glasses, etc. That is, the user device (701) may be an electronic device that connects to a network to receive a data visualization service according to various embodiments of the present disclosure.

[0152] FIG. 8 is a diagram illustrating a data visualization and interaction method according to various embodiments.

[0153] Referring to FIG. 8, the data visualization and interaction method may include a data imaging application step (S810), a data imaging and diagnosis step (S820), a data visualization step (S830), and an interaction step (S840).

[0154] In the data imaging request step (S810), the electronic device may request imaging, diagnosis, or visualization of a data set from the server. Specifically, the electronic device may transmit the data set and request information about the data set to the server. The server may process the data set based on the received data set and request information.

[0155] In the data imaging and diagnosis step (S820), the server can image the data set by vectorizing it. Specifically, the server can identify a vector set corresponding to the data set by embedding the data set in a latent space of a specific dimension.

[0156] The server can diagnose the intrinsic properties and task-dependent attributes of the data set by analyzing the distance, neighbor relationship, or cluster structure between individual data points included in the data set based on the identified vector set.

[0157] For example, to measure data density as an intrinsic characteristic, the server can analyze the distance distribution from a specific data point to the K nearest neighboring vectors, or identify areas overly concentrated in a specific class to determine bias. Furthermore, to assess homogeneity, the server can measure the diversity or dispersion of the region to which each vector belongs.

[0158] Additionally, for example, the server can diagnose task-dependent characteristics by identifying how each class is distributed in vector space (e.g., whether there are boundaries or overlaps between classes) for a data set used in a classification task. This allows the server to detect instances of class imbalance or instances where a particular class is excessively close to another class (i.e., a large number of hard negatives).

[0159] During this process, the server can detect outliers or identify factors that degrade data quality (e.g., incorrect labels, duplicate data, noisy samples, etc.) by considering at least one of the Euclidean distance, cosine similarity, or other distances between vectors. Furthermore, the server can pass these analysis results to subsequent steps (e.g., data cleaning or label correction) to improve the quality of the dataset.

[0160] In the data visualization step (S830), the server can visualize and express the data set in a two-dimensional or three-dimensional space. The server can reduce the dimensionality of the data set or the vector set corresponding to the data set to obtain a data image (IOD) and provide the data image in a two-dimensional or three-dimensional space. For example, the server can apply dimensionality reduction techniques such as PCA (principal component analysis), t-SNE, and UMAP, or use a neural network-based embedding model to map high-dimensional data into a low-dimensional visualization space.

[0161] Here, the server can provide visualized results that reflect the diagnostic results (cluster structure, outliers, bias, class imbalance, etc.) produced in the previous step (e.g., data diagnosis step). Specifically, the server can display data points identified as outliers with a separate color or special marker (e.g., triangles, stars, etc.), or visually highlight areas with ambiguous class boundaries or high-density areas (e.g., region borders, gradients), thereby enabling users to grasp data distribution and quality issues at a glance.

[0162] In the interaction step (S840), the server can process inputs related to the visualized data image received from an electronic device (e.g., a user device) to provide a dynamic response to the visualized data image. Specifically, the user can perform inputs such as touch, mouse click, drag, or pinch zoom on any point within the visualized data image (e.g., a specific data point, cluster, or selected area).

[0163] For example, when a user clicks on a data point, the computing device (server) of the present disclosure may provide a response by retrieving metadata (e.g., individual feature values, label information, statistical values, etc.) associated with the data point and transmitting the information to the user device, which may be displayed in a pop-up window, tooltip, or separate layer. As another example, when the user designates an area by dragging a specific section in the visualized space, the server may further analyze the statistical distribution, density, class distribution, etc. of the data points within the designated area and then provide the results (e.g., mean value, deviation, representative image, number of samples, etc.) to the user device.

[0164] Additionally, users can interact with the visualized data image, such as zooming in, zooming out, rotating (for 3D visualizations), or changing the color scale, to view the data in various ways. In response to these change requests, the server reapplies dimensionality reduction techniques (PCA, t-SNE, UMAP, etc.) or color-shape mapping algorithms, or adjusts parameters to generate an updated data image and deliver it to the user's device.

[0165] In this way, the interaction step (S840) of the present disclosure enables users to intuitively and immediately explore visualized data, allowing them to gain a deeper understanding of the data's inherent structure or to detect potential quality issues (e.g., labeling errors, outliers, unclear boundaries between clusters, etc.) early. Furthermore, this interaction feature can be effectively utilized by machine learning model developers during data cleaning or label correction tasks.

[0166] The various interaction functions described in this section are merely examples to clarify the embodiment, and it is obvious that those skilled in the art can obviously modify and apply response methods to other forms of user input (e.g., gesture input, voice / text commands, etc.) not described above. Therefore, it should be understood that the scope of the present invention is not limited by the specific examples used in this specification.

[0167] FIG. 9 is a drawing for explaining detailed steps performed in an imaging application step according to various embodiments.

[0168] When requesting data imaging and diagnosis, users can input settings for tools to be used for data imaging and diagnosis (e.g., artificial intelligence models for data imaging, such as lenses), data visualization tools, etc.

[0169] Referring to FIG. 9, the data imaging application step (S810) may include a lens level selection step (S811), a lens property determination step (S813), an other setting step (S815), and a visualization tool selection step (S817).

[0170] In the lens level selection step (S811), the user can select a level of a lens including at least one artificial intelligence model to be used for data imaging. In the present disclosure, the level of the lens can be classified and defined according to preset criteria related to the data imaging method. For example, the level of the lens can include, but is not limited to, a first level that diagnoses only quantitative indicators of the data set (basic statistics, missing value ratio, data outlier detection, etc.), a second level that diagnoses by vectorizing the data set using a pre-stored (pre-trained) artificial intelligence model, or a third level that diagnoses by further learning (fine-tuning) or newly learning an artificial intelligence model optimized for the data set and then vectorizing it. For example, when the data set size is very large or high precision is required, the third level can be selected to directly retrain the model to obtain more precise diagnosis results.

[0171] In the lens property determination step (S813), the user can determine the properties of a lens including at least one artificial intelligence model to be used for data imaging. In the present disclosure, the properties of the lens may include the structure of the artificial intelligence model, hyperparameters (e.g., number of layers, number of parameters, learning rate, batch size, etc.), or exploration strategies (e.g., type of optimizer, initialization method). For example, the user can adjust the "model complexity (number of parameters)" or set the "number of learning epochs" to balance analysis and processing time. In addition, the user can also select whether to use a pre-trained model suited to a specific domain (e.g., image, text, structured data, etc.) or a generic model (generic AI model).

[0172] In the other configuration step (S815), detailed options may be provided to reflect user requirements. For example, the user device can set whether to perform data diagnosis by selecting whether to skip the diagnosis process and perform only simple visualization or to perform diagnosis as well. Furthermore, the user device can set whether to visualize the diagnosis results by deciding whether to display cluster density or bias indicators on a separate color scale. Furthermore, the user device can set community creation criteria, such as similarity (distance) criteria between data points or the method of creating subgroups or communities based on specific characteristics (e.g., label information). Furthermore, the user device can set whether to create snapshots that separately store and compare or analyze the data distribution (visualization state) at a certain point in time (viewpoint).

[0173] In the visualization tool selection step (S817), the user can select at least one of multiple pre-linked visualization tools (e.g., PCT, UMAP, T-SNE, etc.). In this case, the user can select the tool based on its characteristics (e.g., dimensionality reduction method, ease of interpretation of results, processing speed, etc.). Furthermore, by selecting a preset visualization dimension (e.g., 2D, 3D), the user can decide whether to view the data in a simple plane (projection) or to interact with it, such as by rotating or zooming in 3D space.

[0174] For example, PCA is fast and intuitive to interpret, while t-SNE and UMAP offer the advantage of more precisely reflecting data clustering structures. Users can choose a visualization tool based on a comprehensive consideration of the characteristics of the dataset, analysis objectives, computing resources, and other factors.

[0175] As described above, in the data imaging application step (S810), users can directly input various settings and decisions to control the data imaging and diagnostic process so that it is optimized for their analysis purposes and environment. This user-centric configuration step facilitates the vectorization, diagnostics, and visualization processes performed in subsequent steps (S820, S830, etc.), ultimately contributing to improved data quality and enhanced AI model performance.

[0176] FIG. 10 is a diagram illustrating the results of data imaging and diagnosis according to the level of the lens and the visualization tool according to various embodiments.

[0177] FIG. 11 is a diagram illustrating a lens builder included in a computing device according to various embodiments.

[0178] The computing device can build a lens for imaging a data set according to settings input by the user, and can diagnose and visualize the data set using the built lens and visualization tool.

[0179] A computing device can build a lens for imaging a data set according to settings input by a user, and can diagnose and visualize the data set using the built lens and visualization tool. Referring to FIG. 10, a computing device (e.g., a server) can provide a data set to a lens builder (LENS BUILDER). Here, the lens builder is a component for building an appropriate lens (e.g., a set of tools for data imaging and diagnosis) by configuring or combining, learning, or tuning tools (e.g., artificial intelligence models, statistical calculators, etc.) for processing the data set, reflecting the user's settings for imaging and diagnosing the data set. The lens builder can be implemented as a separate, independent hardware device or software module, or can be implemented in a form in which at least one processor executes a plurality of instructions stored in memory for building a lens.

[0180] For example, referring to FIG. 11, the lens builder may include, but is not limited to, a model storage unit including a plurality of artificial intelligence models, a model learning unit for learning the artificial intelligence models, an attribute determination unit for determining attributes of the artificial intelligence models, a calculation tool storage unit including a plurality of calculators for calculating statistical characteristics of data, and a lens storage unit for storing constructed lenses.

[0181] Specifically, the model storage unit may include, but is not limited to, a first storage unit storing a plurality of pre-training models (e.g., a first pre-training model, a second pre-training model, 쪋), a second storage unit storing a plurality of foundation models (e.g., a first foundation model, a second foundation model, 쪋), and a third storage unit storing a plurality of adapters (e.g., a first adapter, a second adapter, 쪋). Here, the adapter is a component that is electronically connected to an artificial intelligence model to convert an existing artificial intelligence model, and may be referred to as a data converter, an artificial intelligence converter, a low-rank adapter, or a conversion module. In addition, the calculation tool storage unit may store a plurality of calculators (e.g., a first calculator, a second calculator, 쪋), and the lens storage unit may store a plurality of lenses (e.g., a first lens, a second lens, a third lens, a fourth lens, 쪋), but is not limited thereto.

[0182] Referring again to FIG. 10, the computing device may use a lens builder to construct at least one lens for imaging or diagnosing a data set. At least one lens may include a computational tool storing at least one instruction for processing the data set to derive characteristics associated with the data set.

[0183] For example, a computing device can diagnose a data set by statistically analyzing the data set.

[0184] Specifically, the computing device may generate a first lens (LENS #1) based on a request for a first-level lens. The first lens may include at least one computational tool for diagnosing statistical characteristics of a data set.

[0185] For example, referring to FIGS. 10 and 11, a computing device can use a lens builder to create a first lens including at least one of a plurality of calculators stored in a calculator storage unit.

[0186] For example, the computing device may configure the first lens by selecting or combining at least one of a plurality of calculators (e.g., a missing value detector, an outlier detector, a variance and deviation calculator, a correlation analyzer, etc.) stored in the computational tool storage. For example, if the user requests “basic statistical level diagnosis,” the computing device may activate a calculator that calculates the mean, minimum, maximum, or standard deviation, and if “missing value ratio diagnosis” is additionally selected, the computing device may configure the first lens to include a missing value detector.

[0187] Additionally, the computing device inputs the data set into the first lens and produces a first diagnostic result (1) for the data set through the first lens. st The first diagnostic result can include, for example, basic indicators indicating the status of data quality such as the total number of data, the percentage of missing values, and the percentage of outliers; statistics by specific attribute (mean, variance, skewness, kurtosis, etc.); results of analysis of correlations (Pearson correlation coefficient, Spearman correlation coefficient, etc.); or the number of data by class, bias measures, etc.

[0188] This allows users to quickly understand the overall statistical characteristics of the data set and determine subsequent actions, such as handling missing values ​​or removing outliers.

[0189] As another example, a computing device can perform more in-depth feature diagnosis by vectorizing a data set and analyzing the vectorized data.

[0190] Specifically, the computing device can generate a second lens (LENS #2) using an appropriate pre-trained AI model based on a request for a second level lens.

[0191] For example, referring to FIGS. 10 and 11, the computing device can generate a second lens by loading one or more models that meet user requirements or data formats (images, text, structured data, etc.) from among a plurality of pre-learning models stored in a first storage of the model storage (e.g., a CNN model for computer vision, a Transformer model for text embedding, etc.).

[0192] Additionally, the computing device inputs the data set to at least one pre-trained model, and generates a first vector set (1) corresponding to the data set from at least one layer of the artificial intelligence model. st VECTOR SET) can be obtained. The computing device calculates at least one characteristic value of data included in the data set based on the first vector set, thereby obtaining a second diagnostic result (2) for the data set. nd A second diagnostic result (2nd DIAGNOSIS RESULT) can be obtained. That is, the computing device can analyze the first vector set (e.g., Euclidean distance, clustering, density calculation, etc.) to produce a second diagnostic result (2nd DIAGNOSIS RESULT) that provides higher-level insights such as the inherent characteristics of the data set, class distribution, and bias.

[0193] Second diagnostic results may include, but are not limited to, cluster structure between data in the embedding space (number of clusters, representative center point), density / dispersion between vectors (whether data is excessively concentrated in a specific area), ambiguity of class boundaries (hard negative rate, inter-class distance distribution, etc.), or differences between subgroups based on specific attribute values.

[0194] This allows us to capture the inherent structure or potential bias issues of the data, unlike simple statistical characteristics, and facilitates the identification of the causes of errors or performance degradation that may occur during model learning.

[0195] Additionally, the computing device can learn an artificial intelligence model optimized for the data set based on a request for a third-level lens and then generate a third lens (LENS #3) to perform vectorization using the model.

[0196] For example, referring to FIGS. 10 and 11 , the computing device may generate a third lens including at least one of a plurality of foundation models stored in a second storage unit of the model storage unit or at least one of a plurality of adapters stored in a third storage unit. Specifically, the computing device may generate the third lens by using a model learning unit to learn at least one of the plurality of foundation models or at least one of the plurality of adapters to be optimized for a data set.

[0197] As a specific example, the computing device may generate a third lens by fine-tuning at least one foundation model based on the data set. Alternatively, the computing device may learn at least one adapter using the data set and generate a third lens including the learned adapter and at least one foundation model.

[0198] For example, a computing device can determine a dimension that is optimized for extracting the intrinsic characteristics of a data set, and train at least one artificial intelligence model to vectorize the determined dimension into a latent space. As a specific example, the computing device can calculate various indices for estimating the data distribution (e.g., reconstruction error, cluster quality indices (e.g., silhouette score, Davies-Bouldin Index), classification or regression accuracy, etc.), and analyze how these indices change as the embedding dimension changes. For example, the optimal dimension can be determined by changing the dimension of the latent space to 16, 32, 64, etc. using an autoencoder method and retraining, and searching for the point where the reconstruction error is minimized or the dimension that has the highest performance for a specific task (e.g., classification, clustering), but is not limited thereto.

[0199] Additionally, for example, a computing device can select a candidate model suitable for a specific domain (e.g., medical images, SNS text, etc.) from among large-scale pre-trained foundation models (e.g., Vision Transformer, GPT series models, etc.) and fine-tune it with the user's data set. Alternatively, the foundation model can be fixed and an adapter (e.g., LoRA, Prompt Tuning, etc.) can be trained to update only some of the model parameters, reducing the amount of computation while more precisely tailoring the user's data distribution.

[0200] A third lens that includes a model trained or tuned in this way generates embeddings specific to that dataset, providing higher diagnostic accuracy and insights.

[0201] The computing device inputs the data set into the third lens, embedding the data set into the latent space of optimal dimension, thereby generating a second vector set (2 ndVECTOR SET) can be obtained. In addition, the computing device calculates at least one characteristic value of data included in the data set based on the second vector set, thereby obtaining a third diagnostic result (3) for the data set. rd A third diagnostic result can be obtained. That is, the computing device can analyze the second vector set (e.g., Euclidean distance, clustering, density calculation, etc.) to produce a third diagnostic result that provides higher-level insights into the inherent characteristics of the data set, class distribution, bias, etc.

[0202] The first vector set produced by the second lens (including a pre-trained artificial intelligence model) may be defined based on a predetermined dimension (hereinafter, "the first dimension") determined according to the structure and parameters of the pre-trained model. Conversely, the second vector set produced by the third lens (including an artificial intelligence model optimized for the data set) may be defined based on a dimension (hereinafter, "the second dimension") adaptively determined by analyzing the intrinsic properties of the data set.

[0203] Accordingly, the first dimension along which the first vector set is defined may be different from the second dimension along which the second vector set is defined. The first dimension along which the first vector set is defined may be determined by the artificial intelligence model, and the second dimension along which the second vector set is defined may be adaptively determined depending on the data set. For example, the first dimension according to the second lens may reflect the fixed structure of a specific pre-trained model (e.g., 768 dimensions, 1024 dimensions, etc.), while the second dimension according to the third lens may be optimized by considering the manifold structure of the data set, reconstruction error, or cluster quality indicators.

[0204] Furthermore, according to other embodiments of the present disclosure, since the first and second dimensions are determined differently, the data distributions and characteristics represented by the first and second vector sets may also differ. That is, while the pre-learning model-based embedding using the second-level lens reflects general and universal features, the embedding learned specifically for the data set using the third-level lens can more precisely reflect the specific characteristics of the target data set (bias structure, class boundaries, noise patterns, etc.).

[0205] Accordingly, in a data diagnosis or visualization procedure, the distribution shape, cluster structure, bias index, etc. may differ between the results using the first vector set and the results using the second vector set, and the computing device according to the present disclosure may support a balanced understanding of the unique characteristics and universal features of the data set by comparing and analyzing the two vector sets with each other.

[0206] Additionally, the computing device may provide a vector set corresponding to the data set (e.g., a first vector set or a second vector set) to at least one visualization tool, thereby providing a data image (Image of Data, IOD) visualizing the data set to the user.

[0207] At this time, the computing device can generate and output at least one of a plurality of different types of data images (IODs) depending on the type of visualization tool or the dimension visualized by the visualization tool.

[0208] For example, when using a first visualization tool (e.g., PCA, a 2D-based dimensionality reduction algorithm), the computing device can project the vector set into two dimensions, obtain a two-dimensional first data image (IOD #1), and then provide it.

[0209] Additionally, for example, when using a second visualization tool capable of three-dimensional visualization (e.g., a specific 3D dimensionality reduction algorithm, a WebGL-based 3D viewer, etc.), the computing device can obtain a three-dimensional second data image (IOD #2) and transmit it to a user device (e.g., a PC, a smartphone, an XR device, etc.), thereby enabling the user to check the data distribution from a three-dimensional perspective.

[0210] Additionally, for example, the dimensionality of the data can be reduced and projected into two, three, or other dimensions using a third visualization tool (e.g., t-SNE, UMAP, etc.) to obtain a third data image (IOD #3).

[0211] In this way, the computing device according to the present invention provides a vector set to various visualization tools to generate multiple types of data images, thereby enabling the user to intuitively check the characteristics of data distribution, cluster structure, location of outliers, boundaries between classes, etc. In particular, the user can compare and review different visual representation results (IOD #1, IOD #2, IOD #3, etc.) for the same data set by freely selecting or converting the desired visualization tool and dimension.

[0212] For example, users can initially observe the overall distribution through a two-dimensional visualization, then switch to a three-dimensional visualization to more clearly understand the spatial proximity between specific clusters. Alternatively, users can alternate between applying t-SNE and UMAP to compare the differences between local and global structures, enabling early identification of outliers or biases that could compromise data quality.

[0213] In this way, the "Image of Data (IOD)" in the present disclosure refers to the result of visualizing a vector set, but is not limited to a specific standard (2D / 3D) or a specific technique (PCA / t-SNE / UMAP / Autoencoder-based visualization, etc.). Furthermore, the computing device may receive user interactions (e.g., zooming in and out, clicking on a specific point, specifying a range, changing a color scale, etc.) for the visualization result in real time, and update the visualized data image or display additional details (e.g., metadata corresponding to each point).

[0214] In addition, the computing device can generate and provide a diagnostic report based on the first diagnostic result, the second diagnostic result, or the third diagnostic result. For example, the computing device can store the first diagnostic result (statistical indicators, missing or outlier ratio, etc.), the second diagnostic result (embedding analysis result based on a pre-learning model, cluster distribution, inter-class distance, etc.), or the third diagnostic result (high-precision embedding and feature analysis result obtained through a customized model or adapter) in a database, or manage them separately in a structured format such as JSON or XML, and then combine or arrange them in the form of a dashboard or document to generate a comprehensive report.

[0215] Diagnostic reports may include, but are not limited to, information about data quality ratings, summary information about factors contributing to data quality degradation (e.g., key outliers or sources of bias, indicators of potential label mismatch), recommendations for data improvement (e.g., data diet or bulking recommendations), cluster-specific visualizations (e.g., visual representations of key clusters in embedding space, density or distance measures, etc.), or predictions of learning model performance (e.g., prediction accuracy or F1 Score estimates assuming a particular model).

[0216] These reports can be delivered to users (e.g., data analysts, AI model developers, etc.) in a structured form, and can be immediately utilized in the analysis or model modification stage by being displayed as charts, tables, or heatmaps in a visualization interface (GUI).

[0217] FIG. 12 is a diagram illustrating a method for a computing device to build a data processing model for data imaging and diagnosis based on user input, according to various embodiments.

[0218] Referring to FIG. 12, a computing device or at least one processor included in the computing device may be configured to perform an operation (S1210) of acquiring a first data set. At this time, the computing device may acquire the first data set by receiving the first data set from a user device or by loading a stored first data set.

[0219] In addition, at least one processor may be configured to perform an operation (S1220) of selecting a first level among a plurality of levels classified according to a data processing method based on a user input. Specifically, at least one processor may select the first level based on a user input of selecting the first level among the plurality of levels. In this case, at least one processor may determine a first data processing method corresponding to the first level. At least one processor may determine the first data processing method by determining at least one of tools (e.g., a calculator or an artificial intelligence model, etc.) in which a processing algorithm for processing the first data set is stored.

[0220] In addition, at least one processor may be configured to perform an operation (S1230) of determining at least one attribute for a first data processing model corresponding to the first level. At this time, the first data processing model may include at least one artificial intelligence model. Specifically, the first data processing model may include at least one pre-stored pre-learning model or an artificial intelligence model trained to be optimized for the first data set. For example, based on an input requesting imaging and diagnosis optimized for the first data set, the at least one processor may generate the first data processing model by training the artificial intelligence model optimized for the first data set, but is not limited thereto. In addition, the at least one attribute may include an attribute related to an operation for processing data by the first data processing model. For example, the at least one attribute may include, but is not limited to, the number of parameters or hyperparameters of the artificial intelligence model or graphics card information.

[0221] Additionally, at least one processor may be configured to perform an operation (S1240) of providing, via a first GUI (Graphical User Interface), prior information related to the task of processing the first data set using the first data processing model. In this case, the prior information may be information configured to notify a user regarding the computing device processing the data set and producing results regarding imaging and diagnosis of the data set.

[0222] FIG. 13 is a diagram for explaining information included in dictionary information according to various embodiments.

[0223] Referring to FIG. 13, the prior information may include cost information related to the cost required for processing a data set, result preview information that predicts and displays the results of processing the data set in advance, progress information indicating the progress of processing the data set, reference result information indicating the processing results of cases similar to the data set, or attribute information indicating the properties of the processing model used for processing the data. Accordingly, the user can use the service more smoothly by checking the expected processing cost, predicted results, model progress, etc. in advance at the time of requesting data set imaging and diagnosis.

[0224] In this disclosure, "cost information" may refer to information related to the cost required to process a data set. For example, cost information may include an estimate of the computational cost required to process a data set (e.g., GPU time, CPU core time, memory usage, etc.) or a payment fee (e.g., credits, points, currency, etc.) that a user must pay to use the processing service. Specifically, by presenting information to the user in the form of "estimated computation time of 2 hours, cost of $10 when using 1 GPU," the user can understand in advance the amount of resources and costs consumed based on the size of the data set, model complexity, etc.

[0225] In this disclosure, "result preview information" may refer to information that briefly presents some key results expected after data set processing, allowing the user to estimate the general form or value of the results. In this case, the computing device may briefly display the expected accuracy range after AI model training, a preview of cluster distribution, or a sample visualization image. Furthermore, for example, the result preview information may include data images obtained by simply inputting a data set into a visualization tool. While such data images may not accurately reflect the inherent characteristics of the data set, by providing the visualization results of the data set in advance, they can encourage the user to predict the results.

[0226] This allows users to determine whether the data imaging and diagnostic process of the present disclosure is suitable for their purposes even before obtaining the final results, and to quickly decide to change the direction of the analysis or reset options as needed.

[0227] In this disclosure, "progress information" may refer to information about a processing algorithm or processing status of a data set. For example, the progress information may include the learning progress, training time, or remaining time for optimization learning based on a data set. By checking the current learning stage (initialization, feature extraction, fine-tuning, verification, etc.) and the estimated remaining time in real time, the user can predict the completion time of processing or determine whether additional resources (e.g., computing resources) are allocated. Furthermore, if a problem (e.g., excessive missing values, surge in outliers, etc.) is detected during processing, a warning message may be displayed alongside the progress to provide the user with an opportunity to take action in advance.

[0228] In this disclosure, "reference result information" may refer to information that indirectly estimates expected results or performance indicators by providing examples of processing results previously performed on cases similar to the data set submitted by the user. For example, reference result information may include imaging and diagnostic results data for reference data similar to the input first data set. For example, it may be provided in the form of "When analyzing 100,000 image data of the same category (domain), the average accuracy was 92%, and the processing time was approximately 4 hours."

[0229] This allows users to predict the results of their own data set analysis or establish processing strategies by referencing insights gained from similar data processing cases of third parties.

[0230] In this disclosure, "attribute information" may refer to information about attributes associated with a data processing tool. For example, attribute information may include the number of model parameters or graphics card information. For example, the information may be expressed as, "This processing uses a Transformer architecture-based model (approximately 100 million parameters) and an RTX 3090 GPU." This information allows users to understand model size, resource compatibility (whether internal or cloud GPU), development environment, etc., and predict model utilization and learning performance in advance.

[0231] Referring again to FIG. 12, at least one processor may be configured to perform an operation (S1250) of constructing a first data processing model based on a user input to the first GUI. In this case, the user input to the first GUI may include an approval input for processing the first data set proposed by the computing device.

[0232] That is, the computing device can process the first data set only if, prior to processing the first data set, it provides the user with prior information about the processing of the first data set, and if an approval input is received from the user who has been provided with the prior information.

[0233] FIG. 14 is a diagram illustrating an example of a computing device constructing a first data processing model according to various embodiments.

[0234] Referring to FIG. 14, a computing device or at least one processor included in the computing device may be configured to perform an operation (S1401) of obtaining a plurality of representative values ​​from a plurality of adapters based on a first data set. At this time, the at least one processor may preprocess the first data set to extract a feature value corresponding to the first data set and input the feature value to each of the plurality of adapters. In addition, the at least one processor may obtain a plurality of representative values ​​based on data output from each of the plurality of adapters into which the same data has been input.

[0235] Additionally, each adapter that receives feature values ​​can output a vector set corresponding to the first data set. In this case, at least one processor can obtain a representative value by operating the output vector set in a predetermined manner. The multiple vector sets output from the multiple adapters can be defined based on different dimensions. The multiple adapters can be configured to output vector sets of different dimensions.

[0236] In addition, at least one processor may be configured to perform an operation (S1402) of selecting a first adapter that satisfies a predetermined condition based on a plurality of representative values. At this time, the predetermined condition may include a condition set for determining an adapter optimized for the first data set. Specifically, at least one processor may select the first adapter by identifying at least one adapter corresponding to the lowest (or highest) value among the plurality of representative values. At this time, the first adapter may be implemented to embed the first data set in an optimal dimension representing the first data set. For example, the first adapter may be configured to output a first vector set of a first dimension based on the first data set.

[0237] Additionally, at least one processor may be configured to perform an operation (S1404) of constructing a first data processing model including a first adapter. Specifically, at least one processor may construct a first data processing model including a foundation model or a pre-trained model and the first adapter by communicatively connecting the first adapter to a pre-stored foundation model or a pre-trained model.

[0238] Computing devices can build data processing models optimized for a given dataset by further training or fine-tuning only the adapters based on the dataset. Because computing devices only train adapters on pre-stored AI models, they can minimize computational costs.

[0239] FIG. 15 is a diagram illustrating another example of a computing device constructing a first data processing model according to various embodiments.

[0240] Referring to FIG. 15, the computing device or at least one processor included in the computing device may be set to perform an operation (S1501) of inputting a first data set into a plurality of foundation models.

[0241] In addition, at least one processor may be configured to perform an operation (S1502) of obtaining a plurality of representative values ​​based on data output from a plurality of foundation models. Specifically, each of the plurality of foundation models may output a vector set corresponding to a first data set, and at least one processor may obtain a representative value by operating the output vector set in a predetermined manner. The plurality of vector sets output from the plurality of foundation models may be defined based on different dimensions. The plurality of foundation models may be configured to output vector sets of different dimensions.

[0242] Additionally, at least one processor may be configured to perform an operation (S1503) of determining a first foundation model satisfying a predetermined condition based on a plurality of representative values, and an operation (S1504) of constructing a first data processing model including the first foundation model. Since the specific method of determining the model based on the predetermined condition and the plurality of representative values ​​has been described above, a detailed description thereof will be omitted.

[0243] FIG. 16 is a diagram illustrating another example of a computing device constructing a first data processing model according to various embodiments.

[0244] Referring to FIG. 16, a computing device or at least one processor included in the computing device may perform an operation (S1601) of inputting a first data set into a first foundation model. At this time, the first foundation model may include a plurality of nodes implemented to receive the same input and output different result values. For example, the first foundation model may include at least one hidden layer including a plurality of nodes into which the same input is input in parallel.

[0245] At least one processor may perform an operation (S1602) of obtaining a plurality of representative values ​​from a plurality of nodes included in the first foundation model, wherein the plurality of nodes input feature values ​​for the first data set in parallel. Specifically, the plurality of nodes output a plurality of vector sets based on the feature values ​​for the first data set, and at least one processor may obtain a plurality of representative values ​​from the plurality of vector sets based on a predetermined operation.

[0246] Additionally, at least one processor may be configured to perform an operation (S1603) of activating a first node based on a plurality of representative values. Additionally, at least one processor may be configured to perform an operation (S1604) of constructing a first data processing model including a first foundation model in which the first node is activated.

[0247] Specifically, at least one processor can determine a first node that satisfies a predetermined condition based on a plurality of representative values, and activate the first node so that input is only input to the first node. In other words, at least one processor can select a node optimized for the first data set among the plurality of nodes, and establish a first data processing model by disconnecting communication with nodes other than the selected node.

[0248] FIG. 17 is a diagram illustrating a method of imaging and visualizing a data set using a first data processing model built on a computing device according to various embodiments.

[0249] Referring to FIG. 17, the computing device or at least one processor included in the computing device may input the first data set into the first data processing model after the first data processing model is constructed (S1260). Specifically, when the learning of the artificial intelligence model for vectorization optimized for the first data set is completed, the computing device may determine that the construction of the first data processing model is completed and generate a trigger signal. At least one processor may input the first data set into the first data processing model based on the generated trigger signal.

[0250] Additionally, at least one processor may obtain a first vector set defined in a first-dimensional embedding region from a first data processing model (S1270). Each of the plurality of vectors (or data points) included in the first vector set may correspond to each unit data included in the first data set.

[0251] Additionally, at least one processor may obtain a plurality of characteristic values ​​corresponding to each of a plurality of data included in the first data set based on the first vector set (S1280). Specifically, at least one processor may obtain a plurality of characteristic values ​​corresponding to each of a plurality of data based on a distance value between at least two or more vectors included in the first vector set.

[0252] Additionally, at least one processor may provide a first data image corresponding to the first data set by providing the first vector set to the first visualization tool (S1290). At this time, at least one processor may also recommend a visualization tool appropriate for visualizing the first data set based on the first vector set or a plurality of feature values.

[0253] For example, according to one embodiment of the present disclosure, a computing device may include a visualization tool DB or communicate with a DB server that manages information related to visualization tools in advance, and perform an algorithm for searching and determining an optimal visualization tool based on data characteristics.

[0254] FIG. 18 is a diagram illustrating a method for a computing device to recommend a visualization tool according to various embodiments.

[0255] Referring to FIG. 18, a computing device or at least one processor included in the computing device may perform a data characteristic extraction step (S1810). Specifically, the computing device may analyze or extract characteristic values ​​such as a domain, data type, data capacity, vector distribution density, number per class, etc. for a data set (or vector set). For example, at least one processor may extract statistics such as a data domain, data size (number of samples) and number of dimensions, average distance between vectors, variance, number of clusters, etc., or model learning specifications (required computing resources, processing time, etc.) based on the data set.

[0256] These characteristic values ​​can be used internally as arguments when generating SQL queries or as parameters for performing algorithm logic.

[0257] In addition, the computing device can perform a visualization tool DB query and SQL query generation step (S1820). The computing device (or DB server) can manage metadata for each visualization tool (PCA, T-SNE, UMAP, Autoencoder-based visualization, etc.) in the visualization tool DB (e.g., in table form). At this time, the computing device can generate a SQL query based on a matching rule between data characteristics and tool metadata to search for "candidates that satisfy specific conditions (domain, data size, cluster structure importance, operation time constraints, etc.) among available visualization tools."

[0258] Additionally, the computing device can perform the recommendation algorithm execution and result generation step (S1830). Specifically, when the DB search results are returned in multiple visualization tools, the computing device can apply a score calculation or weight-based algorithm. For example, the computing device can be configured to increase the preference for PCA or UMAP when the amount of data is very large, to increase the score of t-SNE or UMAP when cluster accuracy (detailed clustering) is important, or to prioritize PCA with low computational complexity when real-time interaction is required.

[0259] Based on this, the final recommendation priority is determined, and a list such as "t-SNE (1st), UMAP (2nd), PCA (3rd)" can be presented to the user.

[0260] Additionally, the computing device can perform the final decision step (S1840) on a visualization tool based on a user request. For example, the user may be informed that "t-SNE is the most suitable tool based on data characteristics and priority criteria," and the user can then decide whether to use the recommended tool as is or select a different tool.

[0261] According to one embodiment of the present disclosure, even if a user manually specifies a specific tool, the pre-recommendation results allow the user to recognize key pros and cons (e.g., visualization quality vs. processing time) in advance.

[0262] Additionally, when the finalized visualization tool (e.g., the first visualization tool) is applied, at least one processor can execute the corresponding algorithm on the first vector set to generate and provide a first data image. Optionally, other candidate tools (e.g., the second and third visualization tools) can be sequentially applied to compare or analyze multiple data images.

[0263] In one embodiment of the present disclosure, a computing device can reflect the type of data set (image, text, structured data, etc.), domain (medical, social networking services, finance, etc.), data volume, vector distribution characteristics (density, variance, number of clusters, etc.), processing time constraints, etc., into a "visualization tool recommendation" algorithm, thereby guiding a user to select a visualization tool more rationally. Furthermore, by querying the visualization tool database using SQL to identify available candidates and calculating priorities among the candidates, the user can obtain visualization results that enable efficient understanding of a high-dimensional embedding space.

[0264] In this way, the computing device according to the embodiment of the present disclosure provides an environment in which data distribution or cluster structure can be accurately and intuitively understood by appropriately utilizing the strengths and weaknesses of each technique such as PCA, T-SNE, and UMAP through a process of recommending a visualization tool based on data characteristics.

[0265] FIG. 19 is a diagram illustrating a method for a computing device to generate synthetic data based on data imaging according to various embodiments.

[0266] Referring to FIG. 19, a computing device or at least one processor included in the computing device can generate synthetic data based on a vector set corresponding to a data set.

[0267] Specifically, at least one processor can obtain a vector set by imaging a dataset using a lens including at least one artificial intelligence model, and can visualize the dataset using at least one visualization tool to display a data image (IOD) corresponding to the dataset in a visualization space.

[0268] At this time, at least one processor can generate synthetic data based on the vector set. Specifically, at least one processor can input at least one vector included in the vector set to a calculator (CACULATOR) for which predetermined operation conditions are set. In this case, the calculator can output a latent code based on the at least one vector. In addition, at least one processor can input the output latent code (LATENT CODE) to a generation model (GENERATOR). At this time, the latent code can be any vector defined in the latent space (or embedding space). The generation model can include at least one artificial intelligence model (e.g., diffusion model, GAN, VAE, etc.) for data generation. The generation model can generate and output synthetic data based on the input latent code.

[0269] At least one processor can visualize the generated synthetic data to visually represent the relationship between the synthetic data and the existing data set. Specifically, at least one processor can input the synthetic data into a visualization tool to display a synthetic point (SP) corresponding to the synthetic data on the data image. Alternatively, at least one processor can input the synthetic data into a lens to obtain a synthetic vector (SYNTHETIC VECTOR, not shown) corresponding to the synthetic data, and input the synthetic vector into the visualization tool to display a synthetic point corresponding to the synthetic data on the data image.

[0270] By visually representing the location of synthetic data generated from data images of an existing data set, users can be guided to intuitively recognize the relationship between the generated data and existing data.

[0271] FIG. 20 is a diagram illustrating a method for a computing device to generate synthetic data based on user input according to various embodiments.

[0272] FIG. 21 is a diagram illustrating an example of a method for a computing device to generate synthetic data based on user input, according to various embodiments.

[0273] Referring to FIG. 20, a computing device or at least one processor included in the computing device may visualize a data set and provide a first data image corresponding to the data set (S2010). For example, by utilizing visualization tools such as the aforementioned PCA, t-SNE, and UMAP, each unit data included in the data set may be projected into a two-dimensional or three-dimensional space, and then a data image (IOD) including multiple data points may be displayed on a GUI (Graphical User Interface).

[0274] In addition, at least one processor can receive a user input for a first area on the first data image (S2020). For example, referring to FIG. 21, a user's input such as a click, touch, or drag can be recognized for a first area (R) representing a specific point or range on the first data image (IOD). Here, the user input is an input event that occurs on the GUI, and includes various forms such as a left mouse click, a mobile touch, or a pen drawing.

[0275] Referring again to Figure 20, at least one processor can determine a latent code corresponding to the first region (S2030). The latent code is utilized as a model internal embedding or noise vector when generating synthetic data.

[0276] At this time, at least one processor may request feedback regarding the first region specified by user input. Specifically, the at least one processor may provide the user with a graphical user interface (GUI) (e.g., a pop-up window or highlighting) that visualizes the first region or displays a confirmation message (e.g., "Is this region the target region for generating synthetic data?"). This allows for error prevention by triggering a cancel or reset procedure if the user mistakenly clicks on a point that does not match their intent.

[0277] At this time, at least one processor may define conditions for data generation based on user input. Specifically, at least one processor may define data generation conditions for determining potential code to be input into the data generation model.

[0278] Additionally, at least one processor may determine the latent code in various ways based on the properties of the first region. For example, at least one processor may selectively execute an algorithm that determines the latent code based on the presence or absence of data points or the number of data points contained in the first region.

[0279] For example, if the first region does not contain any data points, at least one processor may determine a latent code based on at least one vector corresponding to at least one point adjacent to the first region. Alternatively, at least one processor may derive a target region including the first region where the user input was made and at least one point adjacent thereto, and then determine a latent code based on at least one vector corresponding to at least one point included in the target region.

[0280] As another example, if the first region includes at least one data point, at least one processor can determine a latent code based on at least one vector corresponding to the at least one data point included in the first region.

[0281] At this time, at least one processor can determine a latent code based on at least two or more vectors. Specifically, at least one processor can identify at least two or more vectors corresponding to at least one point adjacent to the first region, and determine a latent code corresponding to the first region based on the identified at least two or more vectors. For example, by setting a vector interpolated by vectors corresponding to points near the first region as a latent code, synthetic data reflecting the typical (centroid-like) characteristics of the corresponding region can be generated.

[0282] Additionally, if a single point is included in the first region, it can be defined as a target vector, and a latent code can be generated based on the relationship (distance, density, label, etc.) between adjacent vectors around this target vector.

[0283] As another example, if a cluster of multiple data points exists in the first region, at least one processor can determine a potential code based on the properties of the cluster. Specifically, referring to FIG. 21, if the user drag-selects a portion of a specific dense cluster on the IOD, a potential code can be determined based on the centroid or boundary information of the data points belonging to the cluster.

[0284] Referring again to FIG. 20, at least one processor may provide the determined potential code to a generation model to generate synthetic data (S2040). Furthermore, at least one processor may input the generated synthetic data into a data processing model to obtain a synthetic vector corresponding to the synthetic data (S2050). Furthermore, at least one processor may visualize the synthetic vector to provide a synthetic point corresponding to the synthetic data on a first data image (S2060).

[0285] At least one processor may identify at least one area requiring improvement based on diagnostic results for the data set. At least one processor may visually display the identified at least one area via a display of the user device.

[0286] At least one processor can activate at least one identified region to enable interaction. Specifically, at least one processor can enable user input for at least one region requiring improvement. In this case, a data improvement process for the region requiring improvement can be performed based on user input regarding the at least one region requiring improvement.

[0287] At least one processor can extract at least one feature of a plurality of unit data included in the data set based on a vector set that embeds the data set in a latent space of a specific dimension. Based on the at least one extracted feature, the at least one processor can identify data in need of improvement and visually display data points corresponding to the data in need of improvement or a specific region containing the data points.

[0288] Referring to FIG. 21, at least one processor can generate synthetic data by further reflecting a user's prompt input. Specifically, at least one processor can receive a user prompt (PROMPT) through an input interface provided on a GUI (e.g., a text field, voice input, a slider, etc.). The user prompt can express the topic, style, properties, constraints, etc. of the synthetic data in natural language or tag-based. For example, the user prompt can include text instructions in the form of "Generate an image of a cat with a blue background" or "Silhouette format with emphasized perspective." Additionally, if the user prompt specifies a numeric parameter or categorical tag (e.g., "realistic", "cartoon"), the corresponding value can be interpreted as prompt-internal information and passed to the model.

[0289] At least one processor can combine the determined latent code with a user-input prompt to set a condition for a generative model (e.g., a GAN, a VAE, a Diffusion Model, etc.). For example, the latent code can be used as a base latent vector for image generation, and the prompt can be converted into a text embedding (e.g., a CLIP or BERT-based encoder) or a conditional layer and input to the model. This structure can be implemented similarly to how text embeddings condition the overall synthesis process in a "text-to-image" diffusion model (e.g., Stable Diffusion).

[0290] In this case, the generative model (GENERATOR) can determine the basic shape or feature distribution based on the latent code, additionally reflect the subject, style, or detailed properties through prompts (e.g., text, tags, or parameters), and output the final synthetic data. For example, if the latent code is the result of "interpolation between two vectors," it already reflects intermediate properties (e.g., mixing two classes) or spatial characteristics. If the prompt is given as "cat + blue background + cartoon style," the style or thematic elements can be reflected through the internal conditioning path of the model, resulting in a synthetic image with the corresponding characteristics.

[0291] A prompt (PROMPT) entered into an input interface via a GUI along with a determined latent code can be set as a constraint on data generation. At least one processor can input the latent code and the prompt into a generation model, and the generation model can generate synthetic data conditioned by the instructions given by the prompt based on the latent code.

[0292] Additionally, at least one processor can convert the user prompt into a vector form based on a text embedding module (such as CLIP, BERT, GPT) or a conditional network (such as a conditional layer or cross-attention), and then cross-attention it to the hidden layer of the generative model or inject it as a conditional input.

[0293] Additionally, at least one processor may be implemented to reflect user requests at each step (e.g., diffusion step, GAN upsampling step, etc.) based on logic that combines latent codes and text embeddings (e.g., concat, add, attention).

[0294] According to one embodiment illustrated in Fig. 21, after checking the synthesis result, the user can iteratively (refinement synthesis) by modifying the prompt or resetting the latent code (interpolation parameters, selecting additional vectors, etc.).

[0295] According to an embodiment of the present disclosure, customized synthetic data can be easily obtained by immediately reflecting the conditions desired by the user, compared to when only a latent code obtained by simply interpolating multiple vectors is used.

[0296] Additionally, according to the embodiment of the present disclosure, prompt-based generation is advantageous in generating realistic (distribution-preserving) results, since the latent code already reflects the inherent characteristics of the data set.

[0297] Additionally, according to embodiments of the present disclosure, creative uses (e.g., content creation, data augmentation, simulation, etc.) using artificial intelligence models are facilitated by efficiently testing various scenarios (e.g., adversarial examples, artistic expressions, etc.) through prompt changes.

[0298] Additionally, at least one processor can automatically extract or generate a prompt from a user-specified first field, thereby utilizing the prompt as a condition for generating synthetic data. This allows for synthetic results that reflect the inherent properties or metadata of the data, without requiring the user to directly input separate text.

[0299] For example, at least one processor can automatically generate a prompt by analyzing label information, class / category, metadata (e.g., time, location, ID, etc.), visual features (if the data point is an image), etc., for data points belonging to or adjacent to the first region. Specifically, if it is inferred that the region is a "region of cat faces," the at least one processor can generate a simple phrase such as "cat face cluster" or more specific text such as "Multiple cat faces in a close-up shot."

[0300] At least one processor can generate prompts using image captioning technology (e.g., vision-language model, optical character recognition (OCR), etc.), and for structured data, it can automatically convert category names or key attribute names into sentence form.

[0301] As a specific example, at least one processor may use an image captioning model (e.g., a CNN+LSTM architecture, a Transformer-based visual-language model, etc.) to input an image from a first region and automatically output a sentence-based description (caption). The user may be provided with a GUI interface that allows them to "view and edit the extracted phrases," manually modifying the text as desired and then adopting it as the final prompt.

[0302] At least one processor can input the determined latent code or other vector interpolation results and automatically generated prompts into a generative model (GAN, VAE, Diffusion, etc.).

[0303] In this case, the generative model can output conditional synthetic data that reflects the data distribution inferred from the latent code, as well as context, object information, and properties extracted from the user domain.

[0304] In this way, the computing device according to the embodiment of the present disclosure generates linguistic constraints along with latent code, thereby enabling "dimensional expansion" (the generation of data with properties not available with existing data). For example, the computing device can construct an N+3-dimensional data lens by additionally learning synthetic data generated based on linguistic constraints for a data lens that outputs vectors on an N-dimensional latent space.

[0305] According to embodiments of the present disclosure, automatic prompt generation functionality simplifies user input and ensures an intuitive synthetic data generation flow.

[0306] In addition, according to the embodiment of the present disclosure, by converting meta information, visual information, labels, etc. of a user-specified area into “text conditions,” the inherent meaning of the data is naturally reflected.

[0307] In addition, according to the embodiment of the present disclosure, by linking additional models such as captioning or OCR, advanced object recognition or situation description can be performed, thereby enabling flexible application to various domains (e.g., medical imaging, satellite imaging, document processing, etc.).

[0308] Accordingly, according to an embodiment of the present disclosure, a prompt is automatically extracted or generated from a user-specified area and input into a generation model together with a previously determined latent code, thereby obtaining conditional synthetic data reflecting the properties of the area.

[0309] A computing device according to one embodiment of the present disclosure can perform an interaction operation to remove at least some data included in a data set based on a user input.

[0310] FIG. 22 is a diagram illustrating a method for a computing device to remove data based on user input, according to various embodiments.

[0311] Referring to FIG. 22, a computing device or at least one processor included in the computing device may receive a user input regarding a first region including at least one point on a visualized data image (S2210). For example, the first region may be a dense region (a region with very high data density within a cluster) or an region that a user wishes to remove, such as a region containing a large number of outliers. The at least one processor may provide the user with information regarding a region requiring data removal (e.g., a dense region or an outlier region), and the user may provide input regarding the region. The first region may include at least one cluster whose data characteristics satisfy predetermined conditions.

[0312] Additionally, at least one processor may remove at least a portion of the data corresponding to a plurality of points included in the first region (S2220). Here, "at least a portion" refers to only data that corresponds to all or specific conditions (e.g., an outlier filter, a specific label, etc.) specified by the user.

[0313] At least one processor may selectively perform a data removal algorithm based on the properties of the first region. For example, if the at least one processor determines that the data clusters within the first region are excessively dense and cause imbalance, the at least one processor may undersample all data points within the region or remove them at an arbitrary rate. Furthermore, for example, the at least one processor may selectively remove only data that satisfy predefined outlier identification criteria (e.g., distance, density, label mismatch, etc.) among points in the region by determining to remove only outliers within the first region. In this case, the at least one processor may provide refinement options (e.g., filter conditions) to remove only points with a specific label or outside a specific statistical range within a user-specified region (e.g., "Remove only points that fall within this region and have a label of 0").

[0314] Additionally, at least one processor can provide a data image with at least some data removed (S2230). By examining the newly visualized distribution, the user can intuitively understand the effects of improved data imbalance or noise reduction. If necessary, deletion history can be maintained or an Undo function can be provided to prevent data loss due to user error.

[0315] Additionally, at least one processor can perform reimaging and diagnostics (cluster analysis) on the data set after removal to check again how much the data quality has improved and whether there is an advantage in model learning.

[0316] Embodiments of the present disclosure enable direct removal through user interaction, thereby reflecting fine-grained domain knowledge that may be missed by existing automated algorithms.

[0317] In addition, according to the embodiment of the present disclosure, it provides practical advantages such as resolving data imbalance or preventing overfitting during learning through undersampling of dense areas or removal of outliers.

[0318] Additionally, according to embodiments of the present disclosure, immediate feedback can be obtained in a visualized space (2D or 3D), allowing for quick decisions on follow-up actions, such as "how much data has been lost" and "how has the distribution changed".

[0319] Therefore, according to the embodiment of the present disclosure, an intuitive and flexible tool for data quality management can be provided by allowing the user to interactively remove data in a specified area (particularly, a dense cluster area, an outlier-rich area, etc.).

[0320] FIG. 23 is a diagram illustrating a method for a computing device to perform data improvement in response to a data improvement request and provide visual interaction therefor, according to various embodiments.

[0321] Referring to FIG. 23, a computing device or at least one processor included in the computing device may receive a data improvement request (S2310). Specifically, the at least one processor may recognize the data improvement request by receiving user input via at least one GUI that directs data improvement. For example, if a button for improving data in a first manner (e.g., "Resolve data imbalance") or a button for improving data in a second manner (e.g., "Remove noise area") is selected on the user GUI, the at least one processor may recognize that an improvement request has occurred through the corresponding input. Additionally, the user may recognize that a specific class (or attribute) is lacking and request, for example, "Please create more of that class."

[0322] Additionally, at least one processor can visually represent at least one area requiring data improvement (S2320). Specifically, at least one processor can acquire characteristics of the data set based on a vector set corresponding to the data set, and detect areas requiring data improvement based on the characteristics of the data set.

[0323] For example, at least one processor can analyze the density of the data set to identify areas where a specific class is underrepresented or where certain regions (clusters) are overly dense. Furthermore, for example, at least one processor can analyze the bias of the data set to automatically identify areas where the class distribution is imbalanced and requires improvement, such as areas where a specific class is underrepresented or overrepresented. Furthermore, for example, at least one processor can detect areas where a large number of outliers exist (areas where noisy data is concentrated) by performing outlier analysis.

[0324] At least one processor can highlight or outline the detected "areas requiring improvement" on the visualized data image (IOD) to notify the user. In this case, at least one processor can provide interactive guidance (e.g., tooltips) on the GUI, explaining the criteria for selecting the areas. By reviewing this information, the user can intuitively understand the basis for data improvement (e.g., "where is the data lacking, excessive, or noisy?").

[0325] In addition, at least one processor may perform data improvement by generating or removing data for each of at least one region (S2330). For example, at least one processor may generate synthetic data based on the data generation method according to FIG. 19 for a first region on a data image determined to be lacking in data of a specific class. In addition, for example, at least one processor may remove at least some data based on the data removal method according to FIG. 22 for a second region that is over-dense with data or contains outliers. At this time, the user may select at least one of various detailed options (e.g., execute, cancel, adjust detailed options, etc.) for the automatic suggestion.

[0326] Additionally, at least one processor can provide a visualized improved data image corresponding to the results of the data enhancement (S2340). This allows users to visually see how the data distribution has changed, intuitively understand the extent to which imbalances have been resolved, and the extent to which outliers have been reduced. Optionally, at least one processor can also provide users with a comparison view (before / after visualization) with the previous state or a diagnostic report.

[0327] Additionally, users can perform manual corrections by deselecting some of the visualized areas for improvement or by instructing the user to expand the scope of creation or removal. Furthermore, at least one processor can automatically identify areas that require further improvement based on the improvement results, repeatedly performing a loop to improve data quality.

[0328] FIG. 24 is a diagram illustrating an example of a computing device in which a data processing method including a snapshot function is implemented, according to various embodiments.

[0329] Referring to FIG. 24, a computing device (2400) may include multiple components for acquiring snapshot information based on a specific scene on a data image. Here, the multiple components are arbitrarily separated to perform specific operations, and may be physically separate components, or may be separate components based on various operations implemented on a single software program.

[0330] Specifically, the computing device (2400) may include a screener for screening a data image or a vector set corresponding to the data image, an event detector for detecting whether an event for capturing a snapshot has occurred, a capture unit for capturing a snapshot, a correction unit for correcting a captured scene, and a generator for generating various information about the snapshot.

[0331] The screener can screen a vector set or a data image. Specifically, the screener can be configured with at least one metric for measuring characteristic values ​​(e.g., density, bias, homogeneity, presence of outliers, etc.) of the data set based on the vector set or the data image. For example, the screener can analyze basic characteristics (e.g., density, outlier ratio, etc.) in a 2D / 3D data image based on a first metric, or measure high-dimensional distribution characteristics in a high-dimensional vector set based on a second metric.

[0332] An event detector can detect events occurring during screening. Specifically, the event detector can monitor the measured results from the screener and detect an event occurrence when the monitored results satisfy predetermined conditions. At this time, the event detector can store information about the data in which the event occurred (e.g., data items, vector values ​​corresponding to the data, coordinates of data points corresponding to the data, etc.). For example, if the event detector detects a specific area with a higher data density than a threshold value in the screener results, or a specific area with a bias index exceeding a reference value, the event detector can recognize this as an "event occurrence" and generate a signal.

[0333] The generator can generate information about events detected by the event detector. Specifically, the generator can generate tag information by synthesizing event information detected by the event detector, the location and metadata of the captured scene, and feature values ​​calculated by the screener. Alternatively, the generator can record additional information related to the snapshot (class label, time, user ID, etc.). That is, the generator can generate tag information based on information about data in which an event occurred, which is pre-saved or generated from the event detector. The tag information generated by the generator can include, for example, identification labels such as "areas with unusual characteristics (concentrated outliers)" or "sections with excessively high density," the time of snapshot capture, key feature values, etc.

[0334] This allows the computing device to immediately understand the context when the user later retrieves the snapshot.

[0335] The capture unit can capture and save a specific scene on a data image. For example, when a signal notifying the occurrence of an event is transmitted from an event detector or when a user directly commands a snapshot, the capture unit can capture (photograph) a scene containing at least one data point corresponding to an event on the data image at a specific viewpoint. In this case, "capture" includes a process of saving the state of an actual GUI screen (2D / 3D view) or an internal data structure (vector set + visualization mapping parameters) as an image (or video frame).

[0336] The correction unit can perform modifications on captured scenes (snapshots). This provides users with highly visible results, and can also perform graphical corrections, such as visually highlighting areas or adding labels, if necessary.

[0337] For example, a computing device can provide a vector set corresponding to a data set to a screener, and the screener can diagnose the characteristics of the data set based on the vector set. An event detector can identify whether an event has occurred based on the characteristics diagnosed by the screener. For example, if the characteristics diagnosed by the screener satisfy a predetermined condition, the event detector can generate a signal indicating that an event has occurred. In this case, a capture unit can capture a snapshot corresponding to an event based on the signal indicating that an event has occurred. Specifically, the capture unit can generate a snapshot by capturing a scene including at least one data point corresponding to at least one vector associated with the event at a specific point in time. In addition, in parallel with this, a generator can generate tag information based on information about an event that has occurred or information about a location on a data image where an event has occurred. In addition, a correction unit can correct the generated snapshot in a predetermined manner (e.g., brightness adjustment, contrast adjustment, etc.).

[0338] FIG. 25 is a diagram illustrating a method for a computing device to screen a data set to generate snapshot information, according to various embodiments.

[0339] Referring to FIG. 25, a computing device or at least one processor included in the computing device may screen a vector set by calculating at least one characteristic value based on at least one vector included in the vector set using a screener having at least one metric set (S2510). Specifically, at least one processor may diagnose the characteristics of data and measure various indicators such as data density, bias, and homogeneity.

[0340] In the present disclosure, screening may proceed along a specific path based on a specific starting point (coordinate location) on a data image (2D / 3D) or vector set. For example, at least one processor may sequentially scan and analyze a certain range from a predetermined location or a predetermined reference point, thereby evaluating the density, bias, homogeneity, etc. of the data distribution.

[0341] Additionally, at least one processor may perform multiple levels of screening, each level being differentiated based on the screening target. Specifically, at least one processor may perform at least one of a first-level screening for screening a data image or a second-level screening for screening a vector set. The first-level screening performs two-dimensional or three-dimensional data operations based on two-dimensional or three-dimensional data images, while the second-level screening performs high-dimensional operations based on high-dimensional vector sets.

[0342] For example, if a first level of screening is indicated, at least one processor may apply a first metric to measure a characteristic of the data by performing a calculation (two-dimensional or three-dimensional data calculation) based on at least one data point included in the data image.

[0343] Additionally, for example, if a second-level screening is indicated, at least one processor may apply a second metric to measure the characteristics of the data by performing a calculation (high-dimensional calculation) based on at least one vector included in the vector set. For example, at least one processor may diagnose whether the vectors being screened form a structure (clusters, outliers, etc.) on an actual high-dimensional manifold, how high the bias index is, etc.

[0344] Additionally, and not limited thereto, at least one processor may perform multiple levels of screening, each level being differentiated according to the type of characteristic to be screened and diagnosed. Specifically, at least one processor may perform at least one of a first level of screening in which a metric is set for measuring a first type of characteristic (e.g., missing values, data statistics, etc.), a second level of screening in which a metric is set for measuring a second type of characteristic (e.g., density, etc.), or a third level of screening in which a metric is set for measuring a third type of characteristic (e.g., class distribution, etc.).

[0345] Additionally, at least one processor can identify at least one vector whose at least one characteristic value satisfies a predetermined condition (S2520). Through this, at least one processor can identify at least one vector whose first characteristic (e.g., density) satisfies a first condition (e.g., overcrowded area, etc.) or at least one vector whose second characteristic (e.g., bias index) satisfies a second condition (e.g., bias is above a standard).

[0346] Additionally, at least one processor can obtain a data image including a plurality of data points representing the data set in two dimensions or three dimensions by processing the vector set using at least one visualization tool (S2530).

[0347] Additionally, at least one processor can determine a target area including at least one data point corresponding to at least one vector on a data image, and generate tag information associated with the target area (S2540).

[0348] Specifically, at least one processor can identify at least one data point corresponding to at least one vector satisfying a predetermined condition, and determine a target area by determining a predetermined area based on the at least one data point. For example, at least one processor can determine an area having a predetermined radius centered on the at least one data point as the target area.

[0349] Additionally, at least one processor may generate tag information associated with the determined target area. The tag information may reflect diagnostic results associated with the target area. For example, the at least one processor may generate tag information based on information about events detected by the event detector and characteristic values ​​diagnosed by the screener, but is not limited thereto. For example, the tag information may include, but is not limited to, the discovery context (e.g., which metric conditions were met), time, analyst ID, and event type (e.g., unusual section, blank section, overcrowded section, etc.).

[0350] In addition, at least one processor can store a first scene including a target area, the first scene being generated by capturing the target area from a specific viewpoint, and snapshot information including tag information (S2550). Specifically, the at least one processor can acquire the first scene by capturing a view in which the target area is most clearly visible. The at least one processor can acquire the first scene by selecting an optimal scene from among a plurality of captured scenes while adjusting the viewpoint of a virtual camera for capturing the target area. Alternatively, the computing device can automatically calculate a camera angle / magnification that best reveals the target area, adjust the viewpoint, and then take a snapshot (e.g., rotate a 3D point cloud).

[0351] In some cases, at least one processor may compare multiple candidate scenes (e.g., top view, side view, or 45-degree angle view) and provide them to the user, and determine the scene selected by the user as the first scene. In addition, the captured snapshot (first scene) may be corrected (brightness, contrast, highlight, etc.) by the correction unit (MODIFIER) described above, and the final scene corrected in this way and the snapshot information including tag information may be stored in a database or file format.

[0352] Referring again to FIG. 24, the computing device (2400) can generate snapshot information based on user input received via a GUI-based input interface. Specifically, the computing device (2400) transmits a signal to the capture unit instructing the generation of a snapshot based on the user input, and the capture unit can capture at least a portion of the data image (IOD).

[0353] FIG. 26 is a diagram illustrating a method for a computing device to provide snapshot information based on user input according to various embodiments.

[0354] Figure 27 is an example of a screen provided by a computing device.

[0355] Referring to FIG. 26, the computing device or at least one processor included in the computing device can provide a first data image corresponding to a first data set through a first view port (S2610).

[0356] For example, referring to FIG. 27, the computing device may provide a first data image corresponding to a first data set through a first view port (2710). At this time, the computing device may also provide preview information for at least some of the plurality of data points included in the first data image.

[0357] For example, a computing device can provide preview information by outputting the actual data corresponding to a data point. Specifically, the computing device can configure the preview information so that clicking on a specific data point previews the actual data (e.g., original image, text content, statistical values, etc.) corresponding to that point. This allows the user to immediately confirm the meaning of the point in the data image.

[0358] Referring again to FIG. 26, at least one processor may generate first snapshot information including a first scene for a first data image being provided through a first view port in response to a user input received through the first GUI and provide the first snapshot information through a second view port (S2620). Specifically, the at least one processor may generate the first snapshot information by capturing a first scene including at least one data point on the first data image.

[0359] For example, referring to FIG. 27, the computing device may receive user input for a first GUI (2720) and provide first snapshot information including a first scene for a first data image through a second view port (2730). Specifically, the computing device may display the first GUI (2720) and prompt the user to specify a desired scene using a button or a specific gesture (such as dragging or selecting a box) that instructs the user to take a snapshot.

[0360] Additionally, when a user requests a "snapshot," the computing device may transmit a corresponding instruction signal to the capture unit. In this case, the capture unit may capture the first scene by synthesizing the current viewpoint, magnification ratio, or range designation of the first viewport (2710) where the first data image is displayed. Thereafter, the computing device may provide the completed first snapshot information to the user through the second viewport (2730). In the example of FIG. 27, the second viewport (2730) may be used as a "snapshot preview" area.

[0361] Additionally, at least one processor may further provide an interface for receiving memo input from a user via the second viewport. For example, the computing device may activate a "memo input window" (not shown) within the second viewport (2730). While reviewing a snapshot, the user may directly write a string comment (memo) (e.g., "Class A is excessively dense here," "Suspected outlier," etc.). Furthermore, at least one processor may store the user memo by mapping it to the snapshot information, so that the memo information written by the user can be retrieved together when the snapshot is reviewed later.

[0362] Additionally, at least one processor may further provide an interface for generating tag information for the first snapshot information via the second viewport. Specifically, when the first snapshot information is generated by the user, the at least one processor may activate an interface for requesting the user to select at least one of a plurality of tags via at least a portion of the second viewport. In this case, the at least one processor may determine a recommended tag based on the captured scene and provide the user with information about the recommended tag. Additionally, the at least one processor may generate and store tag information based on the selected tag. The at least one processor may structure and store the snapshot information according to the tag type.

[0363] For example, the computing device may display a "tag selection menu" (not shown) in a portion of the second viewport (2730) (e.g., a pop-up, a side panel). At least one processor may analyze the captured scene (e.g., data distribution or point properties within the snapshot) and automatically suggest recommended tags such as "bias," "dense," "outlier," and "label mismatch." When the user selects one or more tags, the processor may generate and store final tag information (e.g., "dense," "bottom-right cluster," "Class A">) based on the selected tags.

[0364] Referring again to FIG. 26, at least one processor may generate link information for connecting the first snapshot information to an external communication network based on user input to the second GUI provided through the second view port (S2630).

[0365] For example, referring to FIG. 27, at least one processor can provide a second GUI (2740) through a second viewport (2720). At least one processor can generate link information based on a user input for the second GUI (2740). The link information can include a link for transmitting the first snapshot information through an external communication network. The link information can be a connection path for transmitting and sharing this snapshot through an external communication network (Internet, company intranet, etc.), and the recipient can check the same scene (snapshot) and tag information or memo information, etc. through the link.

[0366] According to the embodiment of the present disclosure, since the user directly determines and captures an arbitrary point in time / area, it effectively supports customized scenarios (such as enlarging only a specific cluster, emphasizing only a specific label, etc.) that may be missed in automatic capture.

[0367] In addition, by the embodiment of the present disclosure, user comments or classification information are organically combined and stored with snapshots through view port switching (first->second) and memo / tag interfaces, thereby providing richer context for future analysis collaboration or document reporting.

[0368] In addition, according to the embodiment of the present disclosure, the link generation function enables real-time sharing and feedback of snapshot information (scenes, tags, notes, etc.) with external team members or other systems, thereby increasing the efficiency of data interpretation in a remote collaboration environment and inducing external exposure to the solution, which may have an economic ripple effect.

[0369] FIG. 28 is a diagram illustrating a function of a computing device to reproduce snapshot information according to various embodiments.

[0370] A computing device (or at least one processor) according to an embodiment of the present disclosure can sequentially store a plurality of snapshots captured along a screening path and reproduce these snapshots in the form of a video or slide by sequentially providing these snapshots according to a user reproduction instruction.

[0371] Referring to FIG. 28, a computing device or at least one processor included in the computing device may sequentially store a plurality of snapshot information acquired along a screening path (S2810). Specifically, a screener included in the computing device may search a data image (or a high-dimensional vector set) along a specific path (referred to as a "screening path"), and may capture snapshots at each point when a specific characteristic is detected, an event occurs, or at regular intervals. For example, screening may be performed by "inspecting each block (area) on a 2D view while gradually moving from the upper left to the lower right." After checking the characteristic value at each inspection point, a snapshot may be taken if it exceeds a threshold value.

[0372] Additionally, at least one processor can record the captured snapshots in this manner along with connection information (e.g., "Snap #1 -> #2 -> Snap #3 ...") according to the acquisition order or the screening path order. At this time, each snapshot can be stored including metadata such as "view point," shooting location (coordinates along the screening path or high-dimensional mapping information), event cause (excess density, outlier detection, etc.), shooting time," etc.

[0373] This allows snapshots to be stored in a sequence as scanning progresses, allowing the user to visually replay the sequence of points in time later to replay the sequence.

[0374] Additionally, at least one processor can play back multiple sequentially stored snapshot information in a predetermined manner according to a playback instruction input by the user (S2820). Specifically, when a user (e.g., an analyst) presses "Play" or inputs a request to sequentially check screening records through a specific interface, at least one processor can retrieve multiple sequentially stored snapshot information.

[0375] At this time, the computing device may be implemented so that the transmission order of multiple snapshot information during playback differs from the actual snapshot capture order. For example, a user may capture snapshots in a random order at the time of capture, but rearrange them according to specific criteria (e.g., issue priority, reverse chronological order, user-specified order, etc.) during playback.

[0376] Additionally, at least one processor may output snapshots using a predetermined playback method, such as a slideshow method (continuous screen switching at fixed intervals), animation (smooth switching between scenes), or timeline operation (progressing according to the timing of an event in each snapshot). At this time, at least one processor may visually display at least one area where a snapshot was captured on the data image and may also display an animated path connecting the points on the screen. For example, the computing device may provide a dynamic navigation experience to the user by creating a visual effect as if a camera were moving sequentially through the areas where snapshots were captured.

[0377] At this time, the computing device can receive user commands related to playback and control playback based on the received commands. For example, at least one processor can perform actions such as "pause," "skip," "reverse playback," and "add annotations to individual snapshots" based on user interaction actions.

[0378] This allows for a closer look at a specific snapshot during playback, and, if necessary, determines data improvement actions, such as removing or compositing data at that point. Specifically, at least one processor can encourage the user to utilize the improvement feature by providing improvement information (e.g., the reason for the improvement) when an area requiring improvement is identified during snapshot playback.

[0379] A computing device according to an embodiment of the present disclosure may provide a function for linking snapshot information and a diagnostic report. Specifically, the computing device or at least one processor included in the computing device may, during a data screening process, generate a diagnostic report describing the characteristics of a data set along with snapshot information for a specific region.

[0380] FIG. 29 is a diagram illustrating a method for a computing device to generate diagnostic reports and snapshot information in conjunction with each other, according to various embodiments.

[0381] Referring to FIG. 29, a computing device or at least one processor included in the computing device may perform a data screening operation (S2910). Specifically, at least one processor may use a screener to measure specific metrics (such as density, bias, outlier ratio, etc.) for a data set (or vector set) and produce a diagnostic result (e.g., overpopulation in a specific area, class imbalance).

[0382] Additionally, at least one processor can perform a diagnostic report generation operation (S2920). Specifically, at least one processor can generate a diagnostic report by collating or summarizing various diagnostic results. Since the method for generating a diagnostic report has been described above, a detailed description thereof will be omitted.

[0383] Additionally, at least one processor may perform a snapshot information generation operation (S2930). Specifically, if an unusual area (such as an overcrowded area or a cluster of outliers) is discovered during the screening process, at least one processor may capture the scene and generate a snapshot.

[0384] At this time, if detailed diagnostic results are already prepared in report form, at least one processor may record snapshot information in the form of an abbreviated or summarized version of the diagnostic results (e.g., "20 noise points found, clusters with a bias index of 0.85 or higher"). As a specific example, at least one processor may generate snapshot information based on summary information about a specific diagnostic result and a scene captured from a data image corresponding to the diagnostic result.

[0385] Alternatively, when a diagnostic report is generated during the screening process, at least one processor can convert portions of the diagnostic results that meet certain criteria (e.g., above a threshold, a class of interest, etc.) into snapshot information.

[0386] At least one processor can perform a linking operation (S2940) of a diagnostic report and snapshot information. Specifically, at least one processor can link and store the diagnostic result in the diagnostic report and the snapshot information corresponding to the diagnostic result. As a specific example, when at least one processor simultaneously stores two pieces of data (snapshot information, diagnostic report) in a database or file system, the two pieces of data can be linked by linking unique IDs (e.g., report ID, snapshot ID) or specifying identical metadata (e.g., time, area coordinates, event type).

[0387] At least one processor can perform a diagnostic report call operation (S2950) from a snapshot. Specifically, when a user clicks a button instructing to call a diagnostic report in a GUI displaying snapshot information or clicks a label (e.g., class imbalance) displayed on a snapshot, the computing device can quickly call up the diagnostic results corresponding to the snapshot based on the stored information. Thereafter, at least one processor can provide the diagnostic report through a new window (viewport) or pop-up window.

[0388] While reviewing snapshot images, users can load the linked diagnostic report if they're curious about more detailed analysis results. The report can be loaded in a new window (or a secondary viewport) and compared side-by-side with existing snapshots.

[0389] Ultimately, the computing device according to the embodiment of the present disclosure stores snapshot information generated during the data screening process and a diagnostic report detailing the characteristics of the data set in conjunction with each other, and then cross-references them based on user input, thereby supporting the improvement and utilization of data quality by flexibly moving between intuitive visual information and precise analysis results.

[0390] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0391] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. In computing devices, memory; and At least one processor electronically connected to said memory; At least one processor implemented to execute at least one instruction stored in the memory, An operation of providing a first data image corresponding to a first data set, wherein the first data image includes a plurality of data points corresponding to each of the data included in the first data set, through a first view port; An operation of generating first snapshot information including a first scene for the first data image being provided through the first view port in response to a user input received through the first GUI and providing the first snapshot information through the second view port; and A computing device configured to perform an operation of generating link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.

2. In paragraph 1, The actions provided through the above first view port are: A computing device comprising: an operation for providing preview information for at least some of a plurality of data points included in the first data image, wherein the preview information is obtained by outputting actual data corresponding to the data points.

3. In paragraph 1, A computing device characterized in that the first snapshot information is generated by capturing the first scene by considering at least one of a view point, a magnification ratio, or a range specified by a user of the first view port in which the first data image is displayed.

4. In paragraph 1, The actions provided through the above second view port are: A computing device comprising: an interface for receiving a memo input; 5. In paragraph 1, At least one processor, A computing device further configured to perform an operation of activating an interface requesting selection of at least one of a plurality of tags through at least a portion of the second view port when the first snapshot information is generated.

6. In paragraph 1, At least one processor, A computing device further configured to perform an operation of obtaining a corrected scene by performing correction on the first scene.

7. In paragraph 1, A computing device characterized in that the above snapshot information is generated based on a first scene, wherein a scene selected by a user from among a plurality of candidate scenes corresponding to a plurality of viewpoints.

8. In paragraph 1, At least one processor, An operation of generating a diagnostic report by obtaining a diagnostic result for the first data set; and A computing device further configured to perform an operation of linking and storing at least one diagnostic result on the above diagnostic report and snapshot information corresponding to the at least one diagnostic result.

9. In paragraph 8, At least one processor, A computing device further configured to perform an operation of retrieving a diagnostic result included in the diagnostic report corresponding to the first snapshot information based on a user input to the third GUI provided through the second view port.

10. In paragraph 9, At least one processor, A computing device further configured to perform an operation of providing a diagnostic result corresponding to the first snapshot information in the diagnostic report through a third view portlet.

11. In the data interaction method, By at least one processor executing at least one instruction in memory, An operation of providing a first data image corresponding to a first data set, wherein the first data image includes a plurality of data points corresponding to each of the data included in the first data set, through a first view port; An operation of generating first snapshot information including a first scene for the first data image being provided through the first view port in response to a user input received through the first GUI and providing the first snapshot information through the second view port; and A data interaction method comprising: an operation of generating link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port; 12. In paragraph 11, The actions provided through the above first view port are: A data interaction method comprising: providing preview information for at least some of a plurality of data points included in the first data image, wherein the preview information is obtained by outputting actual data corresponding to the data points.

13. In paragraph 11, A data interaction method, characterized in that the above snapshot information is generated based on a first scene, a scene selected by a user from among a plurality of candidate scenes corresponding to a plurality of viewpoints.

14. In paragraph 11, A data interaction method further comprising: when the first snapshot information is generated, activating an interface requesting selection of at least one of a plurality of tags through at least a portion of the second view port; 15. In electronic devices, A display implemented to display at least one GUI (Graphic User Interface); memory; and At least one processor electronically connected to said memory; At least one processor implemented to execute at least one instruction stored in the memory, An operation of providing a first data image corresponding to a first data set, wherein the first data image includes a plurality of data points corresponding to each of the data included in the first data set, through the first view port on the display; In response to a user input received through the first GUI, an operation of providing first snapshot information including a first scene for the first data image being provided through the first view port through the second view port; and An electronic device configured to perform an operation of providing link information for connecting the first snapshot information to an external communication network based on a user input to a second GUI provided through the second view port.

Citation Information

Patent Citations

  • method and device for adjusting an image

    KR1020180051367A

  • Armature Behaviour Improvement type Injector using Dummy Coil

    KR1020250038448A

  • Conveyor Line Can Make Crossing Passage

    KR1020250107605A

  • Method and apparatus for high-dimensional data visualization

    KR102029055B1

  • Method for generating data set

    KR102556766B1