Data clinic method, computer program storing the data clinic method, and computing device for executing the data clinic method
The data clinic method addresses the challenge of evaluating and improving data quality for deep learning models by mapping datasets to embedding spaces and generating improved data images, effectively enhancing data visualization and quality for diverse applications.
Patent Information
- Application Number
- JP2024575535
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-29
- Filing Date
- 2023-06-23
- Publication Date
- 2025-07-23
AI Technical Summary
Existing deep learning models face challenges in accurately evaluating and improving the quality of training data, particularly for unstructured data, as current methods are limited to integrity verification of structured data and lack a comprehensive solution applicable across various technical fields.
A computing device and method for data clinic that involves mapping datasets to an embedding space, identifying point datasets, and generating improved data images through imaging manifolds to enhance data quality and visualization, utilizing algorithms for data imaging, improvement, generation, and evaluation.
Enables efficient and accurate assessment of data quality, allowing for improved data visualization and generation of high-quality virtual data tailored to deep learning models, enhancing their performance across diverse applications.
Smart Images

Figure 2025523515000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a computing device that provides a comprehensive clinical solution for deep learning training data. More specifically, it relates to a program storing a method for improving data by accurately grasping the essential characteristics of a dataset used for training a deep learning model and generating high-quality virtual data, and a computing device for executing the method.
Background Art
[0002] Recently, deep learning-based artificial intelligence algorithms have been utilized in almost all technical fields. In particular, unstructured data without regularity has begun to be used in the field of deep learning, and accordingly, the quantitative problem of data used for learning has emerged.
[0003] The industry has proposed various solutions to solve the quantitative problem of data. In particular, as the technology for generating synthetic data has advanced, synthetic data has been utilized for training deep learning models in various technical fields.
[0004] However, as synthetic data is generated indiscriminately and recently deep learning models based on artificial neural networks have become more advanced, there is a growing need to improve the quality of data rather than further strengthening the quality of the learning model.
[0005] For these reasons, it is important to accurately evaluate the quality of data for training a deep learning model. However, since the limitations of commercially available methods for judging data quality are clear, such as being limited to the integrity verification of structured data, there is a need for a data solution that can be commonly applied to data used in various technical fields.
Summary of the Invention
Problems to be Solved by the Invention
[0006] To solve the above problems, the present disclosure provides a computing device and a data clinic method for a data clinic.
[0007] In addition, the present disclosure provides a method for generating various information for a data set via a computing device and representing it in various ways.
[0008] On the other hand, the problems to be solved in the present disclosure are not limited to the above-mentioned problems, and problems not mentioned can be clearly understood by those having ordinary knowledge in the technical field to which the invention included in the present disclosure belongs from the present specification and the accompanying drawings.
Means for Solving the Problems
[0009] According to an embodiment of the present disclosure, in an operation method of at least one processor included in a computing device, a step of acquiring a data set, a step of confirming a first point data set including respective point data corresponding to respective data included in the data set by mapping the data set to a first embedding space, a step of providing a data image (Image of Data, IOD) acquired by displaying the first point data set in an imaging space, a step of acquiring an improved first point data set including at least one improved point data not included in the first point data set based on the first point data set, and a step of providing an improved data image (Modified Image of Data, MIOD) acquired by displaying the improved first point data set in the imaging space.
[0010] Further, according to an embodiment of the present disclosure, a first converter designed to generate a first manifold based on a dataset defined on an input domain, wherein the first manifold is defined on a first embedding space and includes a first point dataset corresponding to the dataset; a first restorator designed to generate an improved first manifold based on the first manifold, wherein the improved first manifold includes at least one improved point data not included in the first point dataset; and an imaging unit designed to provide a data image by representing the first manifold in an imaging space and provide an improved data image by representing the improved first manifold in the imaging space. A computing device including the above can be provided.
[0011] Further, according to an embodiment of the present disclosure, a computing device for acquiring a dataset and providing information about the dataset, including a memory configured to store a plurality of instructions and at least one processor, wherein the plurality of instructions stored in the memory include a first instruction for instructing an operation of identifying a point dataset obtained by representing the dataset as point data on a latent space based on the dataset, a second instruction for instructing an operation of identifying characteristics of the dataset based on the point dataset, and a third instruction for instructing an operation of providing a data image obtained by representing the point dataset in an imaging space based on the point dataset. The at least one processor can acquire the dataset and selectively execute an operation instructed by at least one or more of the plurality of instructions based on a trigger identified according to the dataset. A computing device including the above can be provided.
[0012] According to an embodiment of the present disclosure, there is provided a computing device for obtaining a dataset and providing a diagnostic result for the dataset, comprising: an output device; a memory; and at least one processor configured to operate based on at least one instruction stored in the memory, wherein the at least one processor is configured to: obtain a dataset; obtain a first manifold by mapping the dataset into a latent space, wherein the first manifold includes a point dataset corresponding to the dataset; obtain a data image by representing at least a part of the point data included in the point dataset in an imaging space; and output, via the output device, a diagnostic report including the data image and additional information obtained by analyzing the data image.
[0013] The means for solving the problems of the present invention is not limited to the above-mentioned means, and the means not mentioned can be clearly understood by those of ordinary skill in the technical field to which the present invention pertains from the present specification and the accompanying drawings. [Advantages of the Invention]
[0014] According to the present disclosure, the inherent characteristics of data can be stored by using a data processing method considering the distribution of data.
[0015] Further, according to the present disclosure, various information about data can be efficiently output by using a data visualization method considering the actual characteristics of data.
[0016] The effects of the present invention are not limited to the above effects, and the effects not mentioned can be clearly understood by those of ordinary skill in the technical field to which the present invention pertains from the present specification and the accompanying drawings.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
BEST MODE FOR CARRYING OUT THE INVENTION
[0018] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. When describing the embodiments, technical content that is well known in the technical field to which the present disclosure pertains and is not directly related to the present disclosure will be omitted. This is to more clearly convey the gist of the present disclosure without obscuring it by omitting unnecessary explanations.
[0019] The embodiments described in this specification are for clearly explaining the spirit of the present invention to those with ordinary knowledge in the technical field to which the present invention pertains. Therefore, the present invention is not limited to the embodiments described in this specification, and the scope of the present invention should be construed to include modifications or variations that do not deviate from the spirit of the present invention.
[0020] The terms used in this specification are general terms that are currently widely used as much as possible considering the functions in the present invention. However, this may change due to the intentions, precedents, or the emergence of new technologies of those with ordinary knowledge in the technical field to which the present invention pertains. However, if specific terms are defined and used in any arbitrary meaning, the meaning of those terms will be described separately. Therefore, the terms used in this specification should be interpreted based not only on the name of the terms but also on the substantial meaning they have and the content throughout this specification.
[0021] The drawings attached to this specification are for easily explaining the present invention, and the shapes shown in the drawings may be exaggerated as necessary to assist in understanding the present invention. Therefore, the present invention is not limited by the drawings.
[0022] In this specification, when it is determined that a specific description of a known configuration or function related to the present invention may obscure the gist of the present invention, the detailed description thereof will be omitted as necessary. Also, the numbers used in the description process of this specification (for example, first, second, etc.) are merely identification symbols for distinguishing one component from another.
[0023] Also, the suffixes "portion" and "part" for the components used in the following description are given or mixed only for ease of preparing the specification, and do not have meanings or roles that are distinct from each other by themselves.
[0024] That is, the embodiments of the present disclosure are provided to make the present disclosure complete and to inform those with ordinary knowledge in the technical field to which the present disclosure pertains of the scope of the present disclosure, and the invention of the present disclosure is defined only by the scope of the claims. Throughout the specification, the same reference numerals refer to the same components.
[0025] Terms such as "first" and / or "second" can be used to describe various components, but the components should not be limited by the terms. The terms are used only for the purpose of distinguishing one component from another. For example, as long as it does not deviate from the scope of the rights according to the concept of the present disclosure, the first component can be named the second component, and similarly the second component can also be named the first component.
[0026] When a component is said to be "connected" or "attached" to another component, it should be understood that it may be directly connected or attached to the other component, or there may be other components in between. On the other hand, when a component is said to be "directly connected" or "directly attached" to another component, it should be understood that there are no other components in between. Other expressions for explaining the relationship between components, that is, "between" and "immediately between" or "adjacent to" and "directly adjacent to", etc. should be interpreted in the same way.
[0027] Each combination of the blocks in the process flowchart diagrams in the drawings and the flowchart diagrams can be executed by computer program instructions. Since these computer program instructions can be loaded onto the processors of general-purpose computers, special-purpose computers, or other programmable data processing apparatuses, the instructions executed via the processors of the computer or other programmable data processing apparatuses generate means for performing the functions described in the flowchart blocks. Since these computer program instructions can also be stored in a computer-usable or computer-readable memory that can be directed to a computer or other programmable data processing apparatus to implement functions in a specific manner, the instructions stored in the computer-usable or computer-readable memory can also manufacture an article of manufacture that includes instruction means for performing the functions described in the flowchart blocks. The computer program instructions may be loaded onto a computer or other programmable data processing apparatus, and thus the instructions that execute a series of operational steps on the computer or other programmable data processing apparatus to generate a process executed by the computer may also be capable of providing steps for performing the functions described in the flowchart blocks.
[0028] Also, each block can represent a module, segment, or portion of code that includes one or more executable instructions for performing a particular logical function. Note also that in some alternative implementations, it is possible for the functions recited in the blocks to occur out of order. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the functions involved.
[0029] As used herein, the term "unit" refers to a hardware component such as software or a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). A "unit" serves a specific role, but is not limited in meaning to software or hardware. It may be configured to be in an addressable storage medium such as a "unit", or may be configured to reproduce more than one processor. Thus, according to some embodiments, a "unit" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided among the components and "units" can be combined with a smaller number of components and "units", or can be further separated into additional components and "units". Further, the components and "units" may be configured to reproduce one or more CPUs within a device or a security multimedia card. Also, according to various embodiments of the present disclosure, a "unit" can include one or more processors.
[0030] Hereinafter, the operating principle of the present disclosure will be described in detail with reference to the accompanying drawings. In the following, when it is determined that a specific description of a known function or configuration related to the description of the present disclosure may obscure the gist of the present disclosure, the detailed description thereof will be omitted. The terms described later are terms defined in consideration of the functions in the present disclosure, and these may vary depending on the intention or convention of the user, operator, etc. Therefore, the definition should be made based on the content throughout this specification.
[0031] The present disclosure relates to a computing device and a system that execute a data clinic method for evaluating the true quality of a dataset for deep learning model learning and providing improvement points.
[0032] FIG. 1 is a diagram for explaining an apparatus and a system for executing a data clinic method according to various embodiments of the present disclosure.
[0033] The data clinic method of the present disclosure can be implemented on a communication network-based platform system 100. Specifically, a server device that collectively processes data, a learning device that learns learning models for various purposes, and a plurality of client devices can be connected to each other on a communication network to transmit and receive data.
[0034] For example, the server device can receive data from at least one of the plurality of client devices, and transmit the received data to the learning device to learn a specific learning model. Further, the server device can process the received data to generate processed data, and transmit the processed data to the plurality of client devices.
[0035] Also, for example, the plurality of client devices can be connected to a server implemented by the server device via a communication network, exchange data with other client devices via the server, or utilize functions implemented by the server.
[0036] Also, the server device, the plurality of client devices, and the learning device can be implemented by one computing device. Specifically, a computing device that executes operations such as learning a deep learning model, collectively processing data, or transmitting and receiving data according to an embodiment can be provided.
[0037] In addition, the server device, the learning device, and the plurality of client devices can include at least one processor (or controller) as at least one computing device.
[0038] Hereinafter, the computing device for providing the data clinic method will be described in more detail.
[0039] FIG. 2 is a diagram showing a block diagram of a computing device that executes a data clinic method and a model learning method for data clinic according to various embodiments of the present disclosure.
[0040] Referring to FIG. 2, the computing device 1000 can include various configurations for providing a data clinic method. Specifically, the computing device 1000 includes a memory 1010 that stores various instructions to be transmitted to data and a processor, a processor 1020 that executes operations based on the instructions transmitted from the memory 1010, and a communication unit 1030 that enables data to be communicated inside the computing device 1000 or enables communication between the computing device 1000 and an external device.
[0041] Also, optionally or alternatively, the computing device 1000 can further include an input device (not shown). At this time, the input device is a device that first receives user input from the outside. For example, the computing device 1000 can further include at least one input device such as a keyboard and a mouse.
[0042] Alternatively, or optionally, the computing device 1000 may further include an output device (not shown). At this time, the output device is a device for displaying specific information externally from the processor 1020. For example, the computing device 1000 may further include at least one output device such as a display, a VR device, AR glasses, an AR projector, or a printing device.
[0043] FIG. 3 is a diagram for explaining various operation methods executed by a computing device that executes a data clinic method according to various embodiments of the present disclosure.
[0044] Referring to FIG. 3, at least one processor 1020 of a computing device for a data clinic can execute various operation methods to execute a data clinic method. At this time, the various operation methods can be encoded and stored in the memory of the computing device. Specifically, at least one processor can process an input input data set based on the various operation methods and output an output data set. At this time, the detailed content of the data included in the input data set and the output data set will be described below (explanation of FIGS. 4 to 33).
[0045] For example, a computing device according to various embodiments of the present disclosure can execute an operation method for data imaging, an operation method for data improvement, an operation method for data generation, an operation method for data feature extraction, or an operation method for data evaluation, but is not limited thereto.
[0046] Also, each of the above operation methods can be executed based on an operation algorithm of at least one processor included in the computing device.
[0047] For example, computing devices according to various embodiments of the present disclosure can execute, but are not limited to, data imaging algorithms, data improvement algorithms, data generation algorithms, data feature extraction algorithms, or data evaluation algorithms, etc.
[0048] At this time, since the names of the respective operation methods and algorithms are arbitrarily named according to the results output for convenience of explanation, each operation method or algorithm is only defined based on the operations executed by the processor, and the names of the operation methods or algorithms themselves do not limit the invention.
[0049] More specifically, a computing device according to various embodiments of the present disclosure can process an input data set according to a data imaging algorithm to generate an image for the input data set.
[0050] Also, a computing device according to various embodiments of the present disclosure can process an input data set according to a data improvement algorithm to improve the data and generate the result of the improvement.
[0051] Also, a computing device according to various embodiments of the present disclosure can process an input data set according to a data generation algorithm to generate synthetic data.
[0052] Also, a computing device according to various embodiments of the present disclosure can process an input data set according to a data feature extraction algorithm to extract the properties of the input data set.
[0053] Also, a computing device according to various embodiments of the present disclosure can process an input data set according to a data evaluation algorithm to evaluate the quality of the input data set.
[0054] Details of each of the above algorithms will be described later.
[0055] Also, computing devices according to various embodiments of the present disclosure can execute the various operation methods or algorithms described above in parallel, sequentially, or selectively. Specifically, the computing device can also use the same input data as input values of different algorithms in parallel, and can also continuously use the result value output according to a specific algorithm as an input value of another algorithm, and can also selectively execute some of the plurality of algorithms according to a preset method.
[0056] Also, the various operation methods or algorithms for the above data clinic can be executed in a deep learning model included in a computing device according to various embodiments of the present disclosure. Specifically, a computing device according to various embodiments of the present disclosure can include one deep learning model for executing the various operation methods or algorithms described above, but is not limited thereto, and may include a plurality of deep learning models for executing each of the above operation methods or algorithms, or may include one or more deep learning models for executing at least a part of the various operation methods or algorithms described above.
[0057] FIG. 4 is a diagram for explaining a method by which a computing device according to various embodiments of the present disclosure provides a data image.
[0058] Referring to FIG. 4, a computing device 1000 according to various embodiments of the present disclosure can receive an input of a data set and provide an Image of Data (IOD).
[0059] At this time, the data set can be M (M>0)-dimensional data. In other words, the data set can be a data set defined on an M-dimensional input space 310.
[0060] Also, the dataset can be a dataset of a single modality. For example, the dataset can be an image dataset. Also, the dataset can be a text dataset.
[0061] Also, without being limited thereto, the dataset can be a set of data with different modalities. For example, the dataset can be an image dataset including annotation information. Also, the dataset can be a mixed dataset of images and text.
[0062] The computing device 1000 according to various embodiments of the present disclosure can input and process, as an input dataset, not only the above-described image data and text data, but also data of all modalities that can be used for deep learning, such as a time-series dataset and a sensor dataset.
[0063] The data image IOD provided by the computing device 1000 according to various embodiments of the present disclosure can process the input dataset and show it in the imaging space 320. Here, the image does not mean a 2D image, but a representation generally referring to a visually represented data. Specifically, the imaging space 320 is a concept including all of a 2D space, a 3D space, and an N-dimensional virtual space, and means a space in which the data image provided by the embodiment appears. For example, when the computing device processes the input dataset and outputs the data image in the PDF format, an output showing the data image in a 2D or 3D imaging space can be output, but it is not limited thereto.
[0064] When the computing device 1000 according to various embodiments of the present disclosure includes an output device (not shown), the computing device can provide a data image via the output device. For example, the computing device 1000 can provide the data image by outputting the data image via a display connected to the computing device 1000. In this case, the imaging space 320 can be the screen of the display. Also, for example, the computing device 1000 can provide the data image by outputting the data image via a printing device connected to the computing device 1000. In this case, the imaging space 320 can be the paper output by the printing device.
[0065] In addition, when the computing device 1000 according to various embodiments of the present disclosure communicates with an external device via a communication unit, the computing device 1000 can provide a data image via the external device. In this case, the imaging space 320 can be the display screen of the external device. For example, when the computing device 1000 is a server device, the server device can provide the data image by transmitting the data image to at least one external device that communicates with the server device via a network connected to the server device.
[0066] The computing device according to various embodiments of the present disclosure can provide a data image including a point data set 330 corresponding to an input data set. At this time, the point data set 330 can be a data set in which each data included in the data set is visualized as a point. In this case, since the shape or color of the visualized points can be variously selected according to the embodiment, the term "point" itself does not limit the invention.
[0067] Also, the points can be expressed by various terms according to embodiments. For example, the points can be expressed by terms such as a vector or a feature that appears in an embedding space or a latent space, but are not limited thereto.
[0068] In order for a computing device according to various embodiments of the present disclosure to provide a data image, as described above, it is necessary to check a point data set corresponding to the input data set.
[0069] At this time, the computing device can obtain the point data set by checking a manifold in which data included in the input data set is formed in an embedding space (or a latent space) of a specific dimension. Here, the manifold can mean a virtual space of a specific dimension in which data actually exists in the dimension of the input space in which the input data set is defined. Also, the manifold can mean any shape in which data is formed in a specific dimension. In other words, the manifold can mean a region in which the point data set is checked or a shape in which the point data set is formed when mapping the input data set to a point data set on an embedding space of a specific dimension.
[0070] Hereinafter, a method in which a computing device according to various embodiments of the present disclosure checks a point data set corresponding to an input data set and provides a data image based thereon will be specifically described.
[0071] A computing device 1000 that provides a data image according to FIG. 4 can include an imaging manifold generation model (not shown) for checking a point data set.
[0072] FIG. 5 is a flowchart for explaining a method by which a computing device according to various embodiments of the present disclosure provides a data image.
[0073] Referring to FIG. 5, a computing device can learn an Imaging Manifold generation model for checking a point data set included in a data image (S1001). At this time, the Imaging Manifold generation model may be a deep learning model including an artificial neural network.
[0074] The computing device 1000 can learn an Imaging Manifold generation model to generate a manifold of a specific dimension in which the intrinsic property of the data set is stored. Here, the intrinsic property of the data set means a property related to the distribution of the data itself, regardless of the modality of the data, the domain in which the data is defined, the category of the data, and the like. For example, the intrinsic property of the data set may include the distance between the data included in the data set. At this time, the distance between the data can mean the Euclidean distance, but is not limited thereto, and can include all mathematical concepts generally used as the distance between data among those skilled in the art.
[0075] The characteristics of the data defined through the present disclosure will be described in more detail below (explanation of FIGS. 10 to 15).
[0076] An example of a method by which a computing device according to various embodiments of the present disclosure learns an Imaging Manifold generation model will be described with reference to FIG. 6.
[0077] FIG. 6 is a diagram showing an example of a method by which a computing device according to various embodiments of the present disclosure learns an Imaging Manifold generation model.
[0078] Referring to FIG. 6, the computing device 1000 can learn an imaging manifold generation model to find a manifold that maintains the unique characteristics of the learning dataset D based on the learning dataset D.
[0079] Also, the computing device 1000 can obtain a first point dataset P1 based on the learning dataset D. At this time, the learning dataset D can be an M-dimensional dataset defined on an M-dimensional input domain.
[0080] Also, the computing device can process the learning dataset D according to preset conditions to obtain the first point dataset P1. Specifically, the computing device can obtain the first point dataset P1 by mapping the learning dataset D to an N-dimensional first embedding space based on a preset condition (for example, a matrix pre-stored for mapping to an embedding space of a specific dimension) defined by a mapping function (f). For example, the computing device can obtain the first point dataset P1 by encoding the learning dataset D, but is not limited thereto.
[0081] A method for determining the optimal dimension of the manifold in which the point dataset is defined will be described in detail with reference to FIG. 9.
[0082] Also, the computing device can obtain a restored dataset D' based on the first point dataset P1. At this time, the restored dataset D' can be an M-dimensional dataset defined on the same M-dimensional space as the learning dataset D.
[0083] In addition, the computing device can process the first point data set P1 according to preset conditions to obtain the restored data set D'. Specifically, the computing device can obtain the restored data set D' by restoring the first point data set P1 to an M-dimensional output domain based on preset conditions (for example, the inverse matrix of a matrix pre-stored for mapping to an embedding space of a specific dimension) defined by the inverse function (f^-1) of the mapping function (f). At this time, the input domain and the output domain may be included in the same virtual space, but are not limited thereto.
[0084] In addition, the computing device can learn an imaging manifold generation model based on the learning data set D and the restored data set D'. Specifically, the computing device can learn the imaging manifold generation model based on a loss function defined based on the similarity between the learning data set D and the restored data set D'. For example, the computing device can learn the imaging manifold generation model in a direction that minimizes the reconstruction error regarding how similar the restored data set D' is restored to the learning data set D, but is not limited thereto.
[0085] FIG. 7 is a diagram for explaining the data imaging process of a computing device according to various embodiments of the present disclosure.
[0086] Referring back to FIG. 5, the computing device can input the dataset into the learned imaging manifold generation model (S1002). At this time, the computing device may receive from the outside that has been learned and input it into the imaging manifold generation model, or a dataset stored in the computing device may be input. For example, the computing device can receive the dataset from an external device connected via a communication network, or call a dataset stored in the memory of the computing device, but is not limited thereto.
[0087] For example, referring to FIG. 7, the computing device 1000 can input a dataset D defined on an M-dimensional input domain into an imaging manifold generation model 700.
[0088] Referring back to FIG. 5 again, the computing device can confirm a point dataset corresponding to the dataset by processing the input dataset with the imaging manifold generation model (S1003).
[0089] For example, referring back to FIG. 7, by outputting a first point dataset P1 defined in an N-dimensional first embedding space based on the input dataset D of the imaging manifold generation model 700, the computing device 1000 can confirm the first point dataset P1. At this time, the first point dataset P1 can form an N-dimensional first manifold.
[0090] In addition, the first point dataset P1 output by the imaging manifold generation model 700 can reflect the relationships between the data included in the dataset D input to the imaging manifold generation model 700. More specifically, as described with reference to FIG. 6, the imaging manifold generation model 700 can be trained to maintain inherent characteristics such as the relevance or similarity between the data included in the input dataset. Thus, when the trained imaging manifold generation model 700 receives an input of the dataset D, it can output a first point dataset P1 that shows the relationships between the data included in the dataset D by generating an N-dimensional manifold. This is because the computing device trains the imaging manifold generation model to minimize the error between the dataset input to the imaging manifold generation model and the dataset restored from the imaging manifold generation model.
[0091] In addition, the first point dataset P1 output by the imaging manifold generation model 700 can correspond to the input dataset D. At this time, each point included in the first point dataset P1 can correspond to each data included in the dataset D. For example, the first image data 711 included in the dataset D can correspond to the first point data 721 included in the point dataset, and the second image data 712 can correspond to the second point data 722.
[0092] In addition, the distance between the points included in the first point dataset P1 output by the imaging manifold generation model 700 can be determined based on the relevance between the data included in the dataset D input to the imaging manifold generation model 700. That is, the higher the relevance (or similarity) between the data included in the dataset D, the closer the data can be located on the first embedding space.
[0093] Further, without being limited thereto, each point included in the point data set can correspond to two or more pieces of data included in the data set. For example, the first image data 711 and the second image data 712 included in the data set can correspond to the first point data 721 included in the point data set.
[0094] Further, without being limited thereto, two or more points included in the point data set can correspond to two or more pieces of data included in the data set. For example, the first image data 711 and the second image data 712 included in the data set can correspond to the first point data 721 and the second point data 722 included in the point data set.
[0095] Also, the computing device 1000 can arbitrarily determine the visual shape of the first manifold defined by the first point data set P1. Specifically, the computing device 1000 can obtain the first point data set P1 by mapping a plurality of point data so that the inherent characteristics of the data set D are maintained in a manifold space having a predetermined shape. For example, the computing device 1000 can store various templates (e.g., spiral shape, etc.) for the shape of the first manifold in advance, and can obtain the first point data set P1 based on at least one of the various templates.
[0096] Also, referring to FIG. 5 again, the computing device can provide the confirmed point data set data image (S1004).
[0097] For example, referring back to FIG. 7, the computing device 1000 can obtain a data image IOD by representing a first point dataset P1 output from the imaging manifold generation model in an imaging domain (or imaging space, 730). Specifically, the computing device 1000 can obtain the data image IOD by mapping the first point dataset P1 from a first embedding space to the imaging space.
[0098] At this time, the computing device 1000 can map the first point dataset P1 according to a preset condition. Specifically, the computing device 1000 can process the first point dataset P1 in a preset manner to obtain the data image IOD.
[0099] For example, the computing device 1000 can represent the first point dataset P1 in the imaging space 730 so that the first point dataset P1 remains as it is, but is not limited thereto.
[0100] Also, for example, the computing device 1000 can generate the data image so that noise data 725 included in the first point dataset P1 is removed. At this time, the noise data 725 can be at least one or more point data located outside the manifold space formed by the first point dataset P1 within the first point dataset P1. In other words, the noise data 725 can be at least one or more outlier points with respect to the manifold space formed by the first point dataset P1.
[0101] The above noise data can be data corresponding to the data included in the dataset. The computing device removing the noise data to provide a data image may be for providing a cleaner image from a visualization perspective.
[0102] An example of a computing device removing noise data from a point data set to provide a data image will be described with reference to FIG. 8.
[0103] FIG. 8 is a flowchart for explaining an example in which a computing device generates a data image based on a point data set according to various embodiments of the present disclosure.
[0104] Referring to FIG. 8, the computing device can confirm a point data set based on the input data set (S1005). Since the technical features of step S1005 have been described above, they will be omitted.
[0105] In addition, the computing device can confirm the manifold region in which the point data set is formed (S1006). At this time, the manifold region can mean a virtual region formed by the point data set in the latent space (or embedding space) defined by the point data set.
[0106] In addition, the computing device can confirm the boundary of the confirmed manifold region (S1007). At this time, the boundary of the manifold region can mean the shape of the manifold region. Specifically, the computing device can determine the boundary of the manifold region connected to the points located on the outer contour from the region where the point data set is located.
[0107] In addition, the computing device can confirm at least one or more point data located outside the boundary of the confirmed manifold region (S1008). Specifically, the computing device can determine at least one or more point data located outside the boundary of the confirmed manifold region as noise data (or outlier data).
[0108] In addition, the computing device can delete the at least one or more point data that has been confirmed (S1009). Specifically, the computing device can enhance the visual effect of the data image by deleting at least one point data determined to be the noise data.
[0109] In addition, the computing device can provide a data image based on the point data set output according to step S1009 (S1010).
[0110] In order to provide a data image that better shows the inherent characteristics of the data set, it is necessary to optimize the manifold formed by the point data set. Here, the optimization of the manifold can mean generating a manifold with minimized reconstruction error through the learning of the manifold generation model described in FIG. 6, but is not limited thereto, and can also mean the process of further optimizing the generated manifold based on another method according to the results of the above learning.
[0111] Specifically, the computing device according to various embodiments can generate a manifold in which the inherent characteristics of the data set are optimized and shown by processing the data set according to a preset method.
[0112] For example, the computing device according to various embodiments can generate a manifold of the point data set so that the noise data is minimized. Specifically, the computing device can repeatedly execute (iterate) the manifold generation process so that the noise data decreases. At this time, the computing device can repeatedly execute (iterate) the manifold generation process until the noise data included in the manifold becomes below a preset standard.
[0113] To provide a data image that better represents the inherent characteristics of a dataset, it is necessary to determine the optimal dimension of the manifold formed by the point dataset. This is because when generating a data image based on a low-dimensional manifold for data processing efficiency, the actual structure of the dataset can be distorted, and when generating a data image based on a high-dimensional manifold for accuracy, the data processing efficiency may decrease.
[0114] To solve the above problems, computing devices according to various embodiments of the present disclosure can determine an optimal dimension for data imaging based on various manifold dimension determination methods.
[0115] Hereinafter, an example of how a computing device according to various embodiments of the present disclosure determines an optimal manifold dimension for data imaging will be described.
[0116] FIG. 9 is a flowchart showing an example of a method for determining an optimal dimension of a manifold for data imaging.
[0117] A computing device can determine an optimal manifold dimension for data imaging based on the minimum reconstruction error according to the dimension of the manifold generated by an imaging manifold generation model. At this time, the minimum reconstruction error may mean the reconstruction error value at the time when the learning of the imaging manifold generation model is completed.
[0118] More specifically, the computing device can generate a manifold by increasing the dimension, and can determine the dimension with the lowest minimum reconstruction error value corresponding to the generated manifold as the optimal dimension. As a specific example, the computing device can determine the dimension at which the minimum reconstruction error no longer decreases as the dimension is increased according to a preset rule as the optimal dimension of the manifold.
[0119] Referring to FIG. 9, the computing device can check the first minimum reconstruction error when generating a first-dimensional manifold (S1011). At this time, the first dimension can be an initial value set for the computing device to execute an algorithm for determining the optimal dimension. For example, when the above algorithm is executed, the computing device can first generate a three-dimensional manifold, but it is not limited thereto.
[0120] In addition, the computing device can check the minimum reconstruction error while increasing the dimension according to a preset rule (S1012).
[0121] At this time, the preset rule can mean a logic for increasing the dimension pre-stored in the computing device. For example, the computing device can check the minimum reconstruction error while increasing the dimension of the manifold by a preset value (for example, 1) at a time, but it is not limited thereto, and can check the minimum reconstruction error while increasing according to a preset sequence (for example, an arithmetic sequence, a geometric sequence, etc.).
[0122] Alternatively, and not limited thereto, the preset rules may be determined based on the first minimum restoration error. More specifically, the computing device can determine the dimension increase width based on whether the first minimum restoration error calculated according to step S1011 is greater than or equal to a threshold value. For example, when the first minimum restoration error is less than or equal to the threshold value, the computing device can increase the dimension by a first increase amount and check the minimum restoration error. When the first minimum restoration error is greater than or equal to the threshold value, the computing device can increase the dimension by a second increase amount greater than the first increase amount and check the minimum restoration error.
[0123] In addition, the computing device can determine the dimension of the manifold as the dimension at which the minimum restoration error no longer decreases (S1013). Specifically, the computing device can determine the dimension value at the point when the minimum restoration error no longer decreases even when the dimension is increased as the dimension of the manifold.
[0124] Alternatively, and not limited thereto, the computing device can determine the dimension of the manifold based on the change amount of the minimum restoration error. Specifically, the computing device can calculate the change amount of the minimum restoration error due to the dimension, and determine the dimension of the manifold by checking whether the change amount of the minimum restoration error is less than or equal to a threshold value. For example, the computing device can determine the dimension value at the point when the change amount of the minimum restoration error is less than or equal to the threshold value as the dimension of the manifold.
[0125] Alternatively, and not limited thereto, the computing device can determine the dimension of the manifold based on the inflection point of the change amount of the minimum restoration error. Specifically, the computing device can determine the dimension value at the point when the change amount of the minimum restoration error increases and begins to decrease as the dimension of the manifold.
[0126] In addition, the computing device can pre-store the maximum dimensional value in which the manifold is defined. Specifically, while the computing device checks the minimum restoration error while increasing the dimension according to the preset rule, when the dimensional value reaches the pre-stored maximum dimensional value, the computing device can determine the pre-stored maximum dimensional value as the dimension of the manifold. At this time, the maximum dimensional value can be set in consideration of the processing capacity of the computing device. This is because the data processing load of the computing device increases as the dimension of the manifold increases, and this is taken into account.
[0127] In addition, the computing device according to various embodiments of the present disclosure can store the dimensional value of the manifold adapted to data imaging according to the input data set. Specifically, the dimensional value of the manifold determined by the above method and the corresponding data set can be pre-stored, and the arbitrarily determined dimensional value of the manifold and the corresponding data set can also be pre-stored. In addition, the computing device can store the relationship between the dimensional value of the manifold and the input data set in the form of a database.
[0128] In addition, the computing device can determine the dimension of the manifold based on a database in relation to the dimension value of the manifold and the input data set. Specifically, when a data set is input, the computing device can check the database for a data set similar to the said data set, and can select the dimension value of the manifold corresponding to the checked data set. For example, the computing device can select the corresponding dimension value by checking in the database for a data set having a distribution similar to the input data set, but is not limited thereto. Also, for example, the computing device can select the corresponding dimension value by checking for a data set having a dimension similar to the input data set, but is not limited thereto. Also, for example, the computing device can select the corresponding dimension value by checking for a data set whose distance from the input data set is equal to or less than a preset threshold value, but is not limited thereto.
[0129] FIG. 10 is a diagram showing a method by which a computing device according to various embodiments of the present disclosure provides characteristics of a data set.
[0130] Referring to FIG. 10, the computing device 1000 can process the acquired data set to obtain the characteristics of the said data set.
[0131] At this time, the "property" of the dataset can mean various information representing the dataset. For example, the property of the dataset can include, but is not limited to, the density, homogeneity, or distribution of the dataset. That is, the property of the dataset can mean an inherent property such as the density of data that has nothing to do with the task for which the dataset is utilized, but is not limited to this, and can also include task-dependent properties such as the ratio of hard-negatives related to the task (e.g., Classification) for which the dataset is utilized.
[0132] Also, the computing device can store in the memory an operation metric corresponding to each of the properties of the dataset. More specifically, the computing device may store, but is not limited to, a metric for computing the density of the dataset, a metric for computing the homogeneity of the dataset, or a metric for computing the distribution of the dataset.
[0133] Also, the computing device can obtain the properties of the dataset based on the stored operation metrics according to a data property extraction algorithm constructed with an artificial neural network. Specifically, the property extraction algorithm can be composed of a feed-forward neural network.
[0134] For example, the computing device can include a separate artificial neural network for computing the properties of the dataset, or can include an artificial neural network including a layer for computing the properties of the dataset, but is not limited to these.
[0135] In one example, when a dataset is input, the computing device can include an artificial neural network for feature extraction designed to extract the features of the dataset. At this time, the artificial neural network for feature extraction can be an artificial neural network that has been transfer - learned to compute the features of the data.
[0136] As another example, the computing device can obtain the characteristics of a dataset by constructing an artificial neural network with a layer added for data feature extraction to an imaging manifold generation model for providing a data image based on the dataset.
[0137] Specifically, the computing device can check a point dataset based on the obtained dataset according to the above - mentioned imaging manifold generation model, and obtain the characteristics of the dataset based on the checked point dataset.
[0138] At this time, the computing device can obtain the characteristics of the dataset by processing each point data included in the point dataset with a preset algorithm. In this case, the computing device can assign a characteristic value to each point data included in the point dataset, and obtain the characteristics of the dataset based on the characteristic value.
[0139] Also, without being limited thereto, the computing device can process the point dataset with a preset algorithm to obtain the characteristics of the dataset.
[0140] FIG. 11 is a flowchart showing a method by which a computing device according to various embodiments of the present disclosure checks the characteristics of a dataset based on the point data included in the point dataset.
[0141] FIG. 12 is a diagram showing an example in which a computing device according to various embodiments of the present disclosure checks the characteristics of a dataset based on a point dataset. The latent space 1250 in FIG. 12 is shown as a two-dimensional space for convenience of explanation, but may actually be a manifold space of three or more dimensions.
[0142] Referring to FIG. 11, the computing device can acquire a point dataset based on a dataset (S1014). At this time, the specific method for the computing device to acquire the point dataset can be directly applied with the above technical features (FIGS. 4 to 9), so it is omitted.
[0143] For example, referring to FIG. 12, the computing device can acquire a point dataset 1200 defined in the latent space 1250 based on the acquired dataset. At this time, the point dataset 1200 can include a plurality of point data including first point data 1201 and second point data 1202.
[0144] In addition, the computing device can calculate a property value for each point data included in the point dataset (S1015). At this time, the property value can mean a value calculated for the point data in order for the computing device to acquire the characteristics of the dataset. Also, the property value can be calculated based on the distance between the point data included in the point dataset. For example, the property value can mean the number of point data existing within a preset distance centered on a specific point data, but is not limited thereto. Also, for example, the property value can mean the average value of the distances to a preset number of point data close to a specific point data, but is not limited thereto.
[0145] Referring back to FIG. 12 again, the computing device can calculate characteristic values according to a preset method for each piece of point data included in the point data set 1200.
[0146] In one example, the computing device can calculate characteristic values based on the number of pieces of point data located in regions 1210 and 1220 within a preset distance centered on a specific piece of point data. For example, the computing device can determine the number of pieces of point data (e.g., 7) located in the first region 1210 within a preset distance from the first piece of point data 1201 as the characteristic value of the first piece of point data 1201. Also, the computing device can determine the number of pieces of point data (e.g., 1) located in the second region 1220 within a preset distance from the second piece of point data 1202 as the characteristic value of the second piece of point data 1202.
[0147] In another example, the computing device can calculate characteristic values based on the average value of the distances to a preset number of pieces of point data adjacent to a specific piece of point data. For example, the computing device can calculate an average distance value based on the distance values to K pieces of point data adjacent to the first piece of point data 1201, and can determine the calculated average distance value as the characteristic value of the first piece of point data 1201, but is not limited thereto.
[0148] As yet another example, the computing device can determine, as a characteristic value of the point data, a class classified for each point data included in the point data set 1200. Specifically, when the data set acquired by the computing device includes annotation information, the computing device can determine the class of each point data included in the point data set 1200 acquired based on the data set. In this case, the computing device can acquire characteristic values based on, but not limited to, the k-NN (k-nearest neighbors) algorithm.
[0149] Referring again to FIG. 11, the computing device can acquire the characteristics of the data set based on the calculated characteristic values (S1016). Specifically, the computing device can acquire the inherent characteristics or task-dependent characteristics of the data set based on the calculated characteristic values. For example, the computing device can acquire, but is not limited to, the density, uniformity, or class distribution of the data set based on the characteristic values of each of the point data.
[0150] Also, the computing device can acquire the characteristics of the data set based on the distribution of the characteristic values of each of the point data. Specifically, the computing device can acquire the characteristics of the data set based on, but is not limited to, the statistical distribution such as the average, deviation, or variance of the characteristic values of each of the point data.
[0151] For example, referring again to FIG. 12, the computing device can determine, as the characteristic of the data set, the average of the characteristic values (for example, the number of point data included in the region within a preset distance) of each of the point data included in the point data set 1200.
[0152] Further, for example, the computing device can determine the statistical distribution of each class of the point data included in the point data set 1200 as a characteristic of the data set.
[0153] FIG. 13 is a flowchart showing a method for a computing device according to various embodiments of the present disclosure to confirm characteristics of a data set based on a point data set.
[0154] FIG. 14 is a diagram showing an example in which a computing device according to various embodiments of the present disclosure confirms characteristics of a data set based on a point data set. The latent space 1450 in FIG. 14 is shown as a two-dimensional space for convenience of explanation, but may actually be a manifold space of three or more dimensions.
[0155] Referring to FIG. 13, the computing device can obtain a point data set based on a data set (S1017). At this time, the specific method for the computing device to obtain the point data set can directly apply the above technical features (FIGS. 4 to 9), so it is omitted.
[0156] For example, referring to FIG. 14, the computing device can obtain a point data set 1400 defined in the latent space 1450 based on the obtained data set.
[0157] Further, the computing device can obtain the characteristics of the data set based on the point data set (S1018). At this time, the computing device can obtain the characteristics of the data set by processing the point data set according to a preset algorithm.
[0158] For example, referring again to FIG. 14, the computing device can obtain the characteristics of the dataset by processing the point dataset 1400 defined on the latent space 1450 according to a preset algorithm.
[0159] Specifically, the computing device can obtain the characteristics of the dataset by processing the point dataset 1400 defined on the latent space 1450 based on the pre-stored filter 1410. At this time, the pre-stored filter 1410 can be a filter with a preset size (for example, a 3×3 or 5×5 kernel).
[0160] In addition, the computing device can apply the pre-stored filter 1410 along a preset path 1420 on the latent space 1450.
[0161] In addition, the computing device can process the point dataset 1400 based on the pre-stored filter 1410 according to the entire latent space 1450 to obtain the characteristics of the dataset.
[0162] In addition, the computing device can process the point dataset 1400 so that the number of point data at the position where the pre-stored filter 1410 is applied in the point dataset 1400 is counted.
[0163] For example, the computing device can obtain the characteristics of the dataset based on the number of point data included in the area where the pre-stored filter 1410 is applied.
[0164] Further, the computing device can obtain the characteristics of the dataset based on the distribution of the number of point data included in the area to which the pre-stored filter 1410 is applied by moving the pre-stored filter 1410 along a preset path 1420.
[0165] Further, when the pre-stored filter 1410 is applied according to the preset path 1420, the computing device can determine the movement range (or stride) of the pre-stored filter 1410. At this time, the movement range of the pre-stored filter 1410 may be predetermined, but is not limited thereto and can be adjusted arbitrarily.
[0166] For example, the computing device can obtain the homogeneity of the dataset based on the deviation or variance (statistical distribution) of the number of point data included in the area to which the pre-stored filter 1410 is applied. In this case, the homogeneity of the dataset can appear as a specific result value based on a lookup table pre-stored in the computing device, but is not limited thereto.
[0167] Further, the computing device can preprocess the point dataset 1400 to obtain information on the positions where the point data exists. In this case, the computing device can apply the pre-stored filter 1410 only to the area corresponding to the position where the point data exists in the latent space 1450.
[0168] Further, the computing device can apply the pre-stored filter 1410 along a preset path 1420 defined on the area corresponding to the position where the point data exists on the latent space 1450.
[0169] As a specific example, the computing device can obtain a feature map related to the characteristics of a dataset by using a convolution algorithm based on a kernel.
[0170] FIG. 15 is a diagram for explaining a method by which a computing device according to various embodiments of the present disclosure obtains the characteristics of a dataset using a convolution algorithm.
[0171] Referring to FIG. 15, the computing device can represent the above point dataset (see reference numeral 1400 in FIG. 14) as a point image 1500 defined by a plurality of eigenvalues. At this time, the plurality of eigenvalues may be values assigned based on whether point data exists at each position on the above latent space (see reference numeral 1450 in FIG. 14). For example, the point image 1500 can be confirmed by representing the position where the point data exists as 1 and the position where the point data does not exist as 0, but is not limited thereto.
[0172] In addition, the size (or dimension) of the point image 1500 can correspond to the size (or dimension) of the above latent space 1450. In FIG. 15, the point image 1500 is shown as a two-dimensional space for convenience of explanation, but may actually be an image of three dimensions or more.
[0173] In addition, the computing device can process the point image 1500 by applying the pre-stored kernel (kernel, 1510) to obtain a feature map (feature map, 1550) related to the characteristics of the dataset.
[0174] Specifically, the computing device can calculate an output value by convolving the point image 1500 based on the pre-stored kernel 1510, and can obtain a feature map 1550 based on the calculated output value.
[0175] At this time, the pre-stored kernel 1510 may be designed to determine the distribution of the dataset. Specifically, it can be a kernel designed to output a feature map 1550 related to the distribution of the input point image 1500 for the pre-stored kernel 1510.
[0176] Therefore, the feature map 1550 can be related to the characteristics of the dataset. For example, the feature map 1550 related to the characteristics of the dataset can be a feature map representing the distribution, density, or homogeneity of the dataset.
[0177] The computing device according to various embodiments of the present disclosure can acquire a dataset and process the acquired dataset so that the dataset is improved. Here, the meaning that the dataset is improved can mean providing a method for improving the quality of the dataset described below (FIGS. 25 to 26), and specifically, it can mean providing a method for improving the dataset in a form adapted by deep learning model learning. For example, the computing device can improve the data by providing a method for making the distribution of the dataset more uniform, but is not limited thereto.
[0178] In one example, the computing device can improve the dataset based on the characteristics of the dataset acquired by the above method.
[0179] FIG. 16 is a diagram for explaining a method by which a computing device according to various embodiments of the present disclosure improves a dataset.
[0180] Referring to FIG. 16, the computing device can check the point data set based on the acquired data set (S1019). At this time, the specific method for the computing device to acquire the data set and check the point data set can directly apply the above technical features (FIGS. 4 to 9), so it is omitted.
[0181] Also, the computing device can check the characteristics of the data set based on the point data set (S1020). At this time, the specific method for the computing device to check the characteristics of the data set can directly apply the above technical features (FIGS. 10 to 15), so it is omitted.
[0182] Also, the computing device can check whether the characteristics of the confirmed data set match the preset criteria (S1021). At this time, the preset criteria can be related to whether the data set needs to be improved. For example, the computing device can check whether the distribution of the data set confirmed based on the point data set matches the preset criteria.
[0183] Also, when the characteristics of the confirmed data set do not match the preset criteria, the computing device can provide an improved point data set so that the characteristics of the data set are adjusted. For example, the computing device can provide the improved point data set by adjusting at least one point data included in the point data set, deleting at least one point data, or adding at least one point data to the point data set, but is not limited thereto.
[0184] A specific example of the computing device providing an improved point data set will be described in more detail with reference to FIGS. 17 and 18.
[0185] FIG. 17 is a diagram showing an example in which a computing device according to various embodiments of the present disclosure generates an improved point data set. The latent space 1750 in FIG. 17 is shown as a two-dimensional space for convenience of explanation, but may actually be a manifold space of three or more dimensions.
[0186] Referring to FIG. 17, the computing device can obtain an improved point data set 1705 by adjusting at least one point data included in the point data set 1700.
[0187] At this time, the computing device can check whether the characteristics of the data set confirmed based on the point data set 1700 match a preset standard. More specifically, the computing device can check whether the characteristics of the data set match a preset standard based on the point data included in at least two or more regions 1710 and 1720 on the latent space 1750 defined by the point data set 1700. At this time, the sizes of the at least two or more regions 1710 and 1720 may all be the same, but are not limited thereto, and may be different from each other. Also, the at least one region 1710 and 1720 may be arbitrarily selected, but are not limited thereto, and may be preset at fixed positions. Also, the at least one region 1710 and 1720 can mean a region to which the filters or kernels of FIGS. 14 and 15 are applied.
[0188] For example, when the difference between the number of point data included in the first region 1710 on the latent space 1750 and the number of point data included in the second region 1720 on the latent space 1750 is equal to or greater than a preset standard, the computing device can adjust at least one point data included in the point data set.
[0189] Specifically, when the difference between the number of point data (e.g., 9) included in the first region 1710 and the number of point data (e.g., 5) included in the second region 1720 is equal to or greater than a threshold value, the computing device can adjust at least one point data included in the point data set 1800 (e.g., adjust the position on the latent space).
[0190] Also, for example, when the difference between the average value of the number of point data included in at least one region 1710, 1720 on the latent space 1750 and the number of point data in a specific region is equal to or greater than a threshold value, the computing device can adjust at least one point data included in the point data set 1800.
[0191] In addition, the computing device can obtain the improved point data set 1705 by adjusting the position of at least one point data included in the point data set 1700 on the latent space 1750. For example, the computing device can obtain the improved point data set 1705 by adjusting the first point 1731 and the second point 1732 defined at the position on the first region 1710 to a specific position on the second region 172.
[0192] In addition, the computing device can determine a position where a point is adjusted on the latent space according to a preset criterion. More specifically, the computing device can determine a position where the first point 1731 and the second point 1732 are adjusted based on the distribution of the point dataset 1700. For example, the computing device can determine a position where the first point 1731 and the second point 1732 are adjusted so that points are uniformly located on the second region 1720. As a specific example, the computing device can move at least one of the first point 1731 and the second point 1732 to an intermediate position between at least two point data with a large distance from each other among the point data included in the second region 1720, but is not limited thereto.
[0193] In addition, the computing device may determine the number of point data to be adjusted so that the distribution of the point dataset 1700 in the point dataset 1700 becomes constant.
[0194] FIG. 18 is a diagram showing another example in which a computing device according to various embodiments of the present disclosure generates an improved point dataset. The latent space 1850 in FIG. 18 is shown as a two-dimensional space for convenience of explanation, but may actually be a manifold space of three or more dimensions.
[0195] Referring to FIG. 18, the computing device can obtain an improved point dataset 1805 by adding at least one point data to the point dataset 1800.
[0196] At this time, the preset criterion for the computing device to generate an improved point dataset may be directly applied to the technical features described in FIG. 17.
[0197] For example, the computing device can obtain an improved point data set 1805 by adding point data to a third region 1810 on the latent space 1850 where the number of point data does not meet a preset criterion. Specifically, the computing device can obtain an improved point data set 1805 by adding a third point 1821 and a fourth point 1822 at any position on the third region 1810.
[0198] In addition, the computing device can determine the position where points are added on the latent space according to a preset criterion. More specifically, the computing device can determine the positions where the third point 1821 and the fourth point 1822 are added based on the distribution of the point data set 1800. For example, the computing device can determine the positions where the third point 1821 and the fourth point 1822 are added so that the points are evenly located on the third region 1810. As a specific example, the computing device can add at least one of the third point 1821 and the fourth point 1822 at an intermediate position between at least two pieces of point data in the point data set 1800 included in the third region 1810 that are far apart from each other, but is not limited thereto.
[0199] In addition, the computing device can determine the number of point data to be added in the point data set 1800 so that the distribution of the point data set 1800 becomes uniform.
[0200] Further, without being limited thereto, the computing device can obtain an improved point data set by removing at least a part of the point data included in the point data set. Specifically, the computing device can obtain an improved point data set by removing at least one point data determined in a preset manner from the point data included in the point data set based on the data set.
[0201] For example, the computing device can remove at least a part of the point data included in the region where the data is excessively dense in the point data set. Specifically, the computing device can select a region including a preset number or more of point data on the manifold region defined by the point data set, and obtain an improved point data set by removing at least one point data included in the selected region.
[0202] As described above, the computing device can improve the data set by adding, adjusting, or removing point data in a direction that corrects the characteristics of the data set determined based on the point data set (or the manifold) to be suitable for the learning of the deep learning model.
[0203] The computing device according to various embodiments of the present disclosure can execute the above data improvement algorithm using a deep learning model. In this specification, the deep learning model for executing the data improvement algorithm is referred to as a "Model for generating Modified Manifold".
[0204] FIG. 19 is a diagram for explaining a method for a computing device according to various embodiments of the present disclosure to learn an improved manifold generation model and provide an improved data image.
[0205] Referring to FIG. 19, a computing device can learn an improved manifold generation model (S1023). The specific method of learning the improved manifold generation model will be described in detail through the description of FIGS. 20 to 22.
[0206] Also, the computing device can confirm a point data set based on the acquired data set (S1024). At this time, the specific method for the computing device to acquire the point data set can be directly applied with the above technical features (FIGS. 4 to 9), so it is omitted.
[0207] Also, the computing device can confirm an improved point data set by inputting the point data set into the improved manifold generation model (S1025). Specifically, the computing device can obtain an improved point data set in which the distance relationship between the point data included in the point data set is adjusted using the improved manifold generation model.
[0208] Also, the computing device can provide an improved data image based on the improved point data set (S1026). The specific method for the computing device to provide an improved data image based on the improved point data set can be directly applied with the above technical features (FIGS. 4 to 9), so it is omitted.
[0209] FIG. 20 is a diagram showing an example of a method for a computing device according to various embodiments of the present disclosure to learn an improved manifold generation model.
[0210] The computing device 1000 can learn an improved manifold generation model to provide a method for improving the acquired dataset into a form adapted by a deep learning model.
[0211] Referring to FIG. 20, the computing device 1000 can check a first point dataset P1 defined in an N-dimensional embedding space based on the acquired dataset. At this time, the specific method of checking the first point dataset can directly apply the above technical features (FIGS. 4 to 9), so it is omitted. At this time, the first point dataset P1 can be checked by defining an N-dimensional first manifold space.
[0212] In addition, the computing device 1000 can check a second point dataset P2 based on the first point dataset P1. At this time, the second point dataset P2 can be checked by defining an L-dimensional second manifold space. Specifically, the computing device 1000 can obtain the second point dataset P2 by representing the first point dataset P1 in an L-dimensional second embedding space.
[0213] In addition, the computing device 1000 can process the first point dataset P1 according to a preset condition to obtain the second point dataset P2. Specifically, the computing device 1000 can map the first point dataset P1 to an L-dimensional second embedding space based on a preset condition (for example, a matrix pre-stored for mapping to an embedding space of a specific dimension) defined by a mapping function to obtain the second point dataset P2. For example, the computing device 1000 can obtain the second point dataset P2 by encoding the first point dataset P1, but is not limited thereto.
[0214] In addition, the computing device 1000 can obtain an improved first point data set P'1 based on the second point data set P2. At this time, the improved first point data set P'1 can be defined on the same N-dimensional first embedding space as the first point data set P1.
[0215] In addition, the computing device 1000 can process the second point data set P2 according to preset conditions to obtain the improved first point data set P'1. Specifically, the computing device can obtain the improved first point data set P'1 by restoring the second point data set P2 to the N-dimensional first embedding space based on preset conditions (for example, the inverse matrix of a matrix pre-stored for mapping to an embedding space of a specific dimension) defined as the inverse function of the mapping function.
[0216] In addition, the computing device 1000 can obtain the improved first point data set P'1 so that the distance relationship between the point data included in the first point data set P1 is adjusted. In other words, the computing device 1000 can learn an improved manifold generation model so that the distance relationship between the point data included in the first point data set P1 is adjusted.
[0217] In addition, the computing device 1000 can adjust the distance relationship between the point data so that the distribution of the first point data set P1 is improved. More specifically, the computing device 1000 can adjust the distance relationship between the point data by moving the point data located in the region with a high density of point data to the region with a low density of point data in the first point data set P1.
[0218] Further, the computing device 1000 can learn the improved manifold generation model based on a loss function defined based on the distances of the point data included in the first point data set P1. For example, the computing device 1000 can learn the improved manifold generation model to extract at least one pair of point data whose distance relationship needs to be adjusted among the point data included in the first point data set P1. Also, for example, the computing device 1000 can learn the improved manifold generation model to add (or synthesize) point data to an area where the distance relationship needs to be adjusted in the first point data set P1.
[0219] FIG. 21 is a flowchart for explaining an example of a method by which a computing device according to various embodiments of the present disclosure learns an improved manifold generation model.
[0220] FIG. 22 is a diagram showing a method by which a computing device according to various embodiments of the present disclosure learns an improved manifold generation model by extracting hard negative pairs. The latent space 2250 in FIG. 22 is shown as a two-dimensional space for convenience of explanation, but may actually be a manifold space of three or more dimensions.
[0221] Referring to FIG. 21, the computing device can perform initial clustering based on the first point data set (S1027). At this time, the initial clustering means clustering a plurality of point data included in the first point data set into at least one or more groups. Specifically, the computing device can cluster the plurality of point data included in the first point data set into at least one or more groups based on the similarity of the data corresponding to the plurality of point data.
[0222] In addition, the computing device can perform initial clustering based on the similarity information for the first point data set. At this time, the computing device may obtain the similarity information from the outside or generate the similarity information in order to obtain the similarity information.
[0223] In one example, the computing device can receive similarity information for the first point data set from the outside. Specifically, the computing device can receive information from the user regarding the similarity of at least two or more point data included in the first point data set. That is, the user can input whether at least two or more point data are similar in the first point data set confirmed from the computing device. For example, the computing device can cluster the first point data set into at least one or more groups based on annotation information for the data set received from the outside, but is not limited thereto.
[0224] In another example, the computing device can obtain similarity information for the first point data set through unsupervised learning. Specifically, the computing device can cluster the first point data set into at least one or more groups by learning the similarity between the point data included in the first point data set by itself. In addition, the similarity information for the first point data set can be confirmed based on the characteristic values of the point data included in the first point data set. Specifically, the computing device can determine that the higher the similarity between the characteristic values of the point data, the higher the similarity.
[0225] As a specific example, referring to FIG. 22, the computing device can perform initial clustering based on the first point data set 2200. Specifically, the computing device can cluster the first point data set 2200 into a first group including the first point data 2215 and a second group including the second point data 2225. At this time, the point data included in the same group can have similar features (positive) to each other. Also, the point data included in the first group and the point data included in the second group can have different features (negative) from each other. For example, the point data included in the same group can be data on the latent space 2250 that can derive similar results when performing a specific task, but is not limited thereto. In FIG. 22, the point data included in the first group is represented by circular points, and the point data included in the second group is represented by square points, but this is only an exemplary representation and does not limit the invention to the representation in the drawings.
[0226] In the present disclosure, a pair of point data clustered into different groups is defined as a negative pair, and a pair of point data clustered into the same group is defined as a positive pair.
[0227] Also, referring to FIG. 21 again, the computing device can perform hard negative pair mining (S1028) based on the initially clustered first point data set. Here, the hard negative pair means a negative pair among the above negative pairs that are close to each other in distance and thus difficult to distinguish from each other.
[0228] Specifically, the computing device can extract the hard negative pairs based on the distance relationship of the clustered first point data set.
[0229] As an example, when there is negative point data within a group different from specific point data and the distance from the specific point data is less than or equal to a threshold, the computing device can determine the specific point data and the negative point data as a hard negative pair.
[0230] As another example, for specific point data, when negative point data in another group is located closer than positive point data in the same group, the computing device can determine the specific point data and the negative point data as a hard negative pair.
[0231] In addition, the computing device can extract positive pairs that are located far from each other in the latent space even though they are in the same group.
[0232] As an example, when there is positive point data within a group the same as specific point data and the distance from the specific point data is greater than or equal to a threshold, the computing device can determine the specific point data and the positive point data as a positive pair.
[0233] As another example, for specific point data, when negative point data in another group is located closer than positive point data in the same group, the computing device can determine the specific point data and the positive point data as a positive pair.
[0234] As a specific example, referring to FIG. 22 again, the computing device can extract hard negative pairs and positive pairs based on the similarity and distance relationships of the point data included in the first point data set 2200 defined on the latent space 2250.
[0235] Specifically, the computing device can determine the reference point data 2201 and the first point data 2215 as a hard negative pair 2210 by checking the first point data 2215 that is included in a group different from the reference point data 2201 but satisfies a preset distance condition for extracting hard negative pairs. Also, the computing device can determine the reference point data 2201 and the second point data 2225 as a positive pair 2220 by checking the second point data 2225 that is included in a group different from the reference point data 2201 but satisfies a preset distance condition for extracting positive pairs.
[0236] Also, referring to FIG. 21 again, the computing device can obtain an improved first point data set by adjusting the distance between the extracted hard negative pairs (S1029). Specifically, the computing device can adjust the position of at least one point data included in the first point data set on the latent space so that the hard negative pair becomes an easy negative pair. At this time, the easy negative pair means a negative pair among the above negative pairs that are far apart from each other and are easy to distinguish from each other.
[0237] In addition, the computing device can obtain an improved first point data set by adjusting the distance between the extracted positive pairs. Specifically, the computing device can adjust the position of at least one point data in the latent space of the first point data set so that the distance between the positive pairs is smaller than a preset distance.
[0238] As a specific example, referring to FIG. 22 again, the computing device can obtain an improved first point data set 2205 by adjusting the position of the first point data 2215 on the latent space 2250 that was identified as a hard negative pair with respect to the reference point data 2201.
[0239] In addition, the computing device can obtain an improved first point data set 2205 by adjusting the position of the second point data 2225 on the latent space 2250 that was identified as a positive pair with respect to the reference point data 2201.
[0240] Also, without being limited thereto, steps S1028 and S1029 in FIG. 21 may be replaced with the following operations.
[0241] For example, the computing device can obtain an improved first point data set by adding at least one point data to the initially clustered first point data set. In this case, the computing device can adjust the distance relationship between the point data included in the first point data set by adding the at least one point data.
[0242] As a specific example, the computing device can determine an area of interest that needs to adjust the distance relationship based on the initially clustered first point data set. At this time, the area of interest may be an area including the above hard negative pairs. In this case, the computing device can adjust the distance relationship between the area of interest and the point data located around the area of interest by generating at least one point data for at least a part of the area of interest. For example, the computing device can adjust the distance between the hard negative pairs to be farther by generating at least one point data in the area between the hard negative pairs on the area of interest including the hard negative pairs.
[0243] FIG. 23 is a diagram showing an operation in which a computing device according to various embodiments of the present disclosure provides an improved data set including virtual data based on a data set.
[0244] FIG. 24 is a diagram showing an example of an operation in which a computing device according to various embodiments of the present disclosure provides an improved data set including virtual data based on a data set.
[0245] Referring to FIG. 23, the computing device can check a point data set based on the acquired data set (S1030). Further, the computing device can check an improved point data set based on the confirmed point data set (S1031). At this time, since the operations S1030 and S1031 can directly apply the above technical features (FIGS. 16 to 22), detailed descriptions thereof are omitted.
[0246] As a specific example, referring to FIG. 24, the computing device 1000 can check the point data set 2410 based on the acquired data set 2400. At this time, the computing device 1000 can obtain the point data set 2410 by mapping the data set 2400 to a latent space based on a preset mapping function. At this time, the point data set 2410 can include first point data 2411 and second point data 2412. For example, the first point data 2411 and the second point data 2412 may be data clustered into different groups, but are not limited thereto, and may be unclustered or data clustered into the same group.
[0247] Further, the computing device 1000 can obtain an improved point data set 2420 based on the point data set 2410. At this time, the computing device can obtain the improved point data set 2420 by processing the point data set 2410 based on a pre-stored improvement algorithm 2430. Specifically, the computing device 1000 can map the point data set 2410 to another latent space according to preset conditions and then restore it to the latent space to obtain the improved point data set 2420. Further, the improved point data set 2420 can include improved first point data (modified first point data, 2421) and improved second point data 2422. For example, the improved first point data 2421 can be obtained by adjusting the position of the first point data 2411 in the latent space, and the improved second point data 2422 can be obtained by adjusting the position of the second point data 2412 in the latent space. That is, the first point data 2411 can correspond to the improved first point data 2421, and the second point data 2412 can correspond to the improved second point data 2422. Further, the improved point data set 2420 can further include third point data 2423. At this time, the third point data 2423 can be point data not included in the point data set 2410. In other words, the computing device 1000 can generate arbitrary third point data 2423 based on the improvement algorithm 2430. That is, the point data set 2410 may not include point data corresponding to the third point data 2423.
[0248] Referring back to FIG. 23 again, the computing device can obtain synthetic data based on the improved point data set (S1032). At this time, the synthetic data can mean data arbitrarily generated by the computing device according to a preset algorithm. Specifically, the synthetic data can be data having the same modality as the acquired data set, but data not included in the data set. More specifically, the computing device can process the improved point data set based on a preset algorithm to generate the synthetic data.
[0249] Also, the computing device can provide a modified data set including the synthetic data (S1033). At this time, the modified data set can include at least one data not included in the data set.
[0250] As a specific example, referring back to FIG. 24 again, the computing device 1000 can provide a modified data set 2450 based on the improved point data set 2420. At this time, the computing device 1000 can provide a modified data set 2450 including at least one synthetic data by generating at least one synthetic data based on the modified data set 2450.
[0251] Further, the computing device 1000 can provide the improved data set 2450 by restoring the improved point data set 2420 to the output domain using the inverse function of the mapping function used to obtain the point data set 2410. Specifically, the computing device 1000 can obtain virtual data by restoring the point data included in the improved point data set 2420, and can provide the improved data set 2450 including the virtual data.
[0252] Also, each data included in the improved data set 2450 can correspond to each point data included in the improved point data set 2420. For example, the computing device 1000 can obtain the first virtual data 2451 based on the improved first point data 2421, obtain the second virtual data 2452 based on the improved second point data 2422, and obtain the third virtual data 2453 based on the third point data 2423. That is, the first virtual data 2451 can correspond to the improved first point data 2421, the second virtual data 2452 can correspond to the improved second point data 2422, and the third virtual data 2453 can correspond to the third point data 2423.
[0253] Also, the improved data set 2450 may include at least one data not included in the data set 2400. Also, the improved data set 2450 may not include at least one data included in the data set 2400. Also, the number of data included in the improved data set 2450 can be equal to or greater than the number of data included in the data set 2400.
[0254] As described above, the computing device can generate virtual data in a neural rendering method based on data improvement, but is not limited thereto.
[0255] Computing devices according to various embodiments can generate virtual data in a CG-based rendering method based on the data improvement. Specifically, the computing device can generate virtual data by generating CG parameters based on the generated improved point data set. More specifically, the computing device can generate virtual data by obtaining a rendering parameter based on at least one point data included in the improved point data set. For example, the computing device can generate the virtual data by implementing the inverse function of the mapping function in a CG rendering model, but is not limited thereto.
[0256] FIG. 25 is a diagram showing an operation of providing the quality of a data set acquired by a computing device according to various embodiments of the present disclosure.
[0257] Referring to FIG. 25, the computing device can acquire a data set (S1034). Also, the computing device can acquire a data image based on the acquired data set (S1035). At this time, since the operation S1035 can directly apply the above technical features (FIGS. 4 to 9), a detailed description thereof is omitted. Also, the computing device can acquire the characteristics of the data set based on the data set (S1036). At this time, since the operation S1036 can directly apply the above technical features (FIGS. 10 to 15), it is omitted.
[0258] In addition, the computing device can provide the quality of the data set based on at least one of the characteristics of the data image and the data set (S1037). Specifically, the computing device can obtain at least one index based on at least one of the characteristics of the data image and the data set, and can provide the quality of the data set based on the at least one index. For example, the computing device can provide the quality of the data set based on an index including the "appropriateness of distribution", "learning fitness", "similarity between data", or "appropriateness of the number of data" of the data set. At this time, the computing device can evaluate the at least one index with various grades, and can provide the final quality for the data set based on the grades assigned to each index.
[0259] In one example, the computing device can evaluate the "appropriateness of distribution" based on the characteristics of the data image or the data set. At this time, the "appropriateness of distribution" can mean how uniform the distribution of the data set is. More specifically, the computing device can evaluate the "appropriateness of distribution" based on the uniformity of the data distribution appearing on the data image or the density (or uniformity) of the data set included in the characteristics of the data set. For example, when the distribution of the data is uniform, the computing device can evaluate the grade of the "appropriateness of distribution" of the data set as "Great", but is not limited thereto.
[0260] In another example, the computing device can evaluate the "learning fitness" based on the characteristics of the data image or the data set. At this time, the "learning fitness" can mean how well the data set fits the learning of a specific deep learning model. More specifically, the computing device can evaluate the "learning fitness" based on the task-dependent characteristics included in the characteristics of the data set. For example, the computing device can evaluate whether the data set fits the learning of an image classification model by determining whether the data corresponding to the class to be classified by the data set is evenly included, but is not limited thereto.
[0261] In yet another example, the computing device can evaluate the "similarity between data" based on the characteristics of the data image or the data set. At this time, the "similarity between data" can mean how similar the data included in the data set is. More specifically, the computing device can evaluate the "similarity between data" based on the distance in the latent space between the data included in the data set.
[0262] As yet another example, the computing device can evaluate the "appropriateness of the number of data" based on the characteristics of the data image or the data set. More specifically, the computing device can evaluate whether the data set includes an appropriate number of data for learning a deep learning model.
[0263] FIG. 26 is a diagram showing an operation of providing an achievable quality of a data set acquired by a computing device according to various embodiments of the present disclosure.
[0264] Referring to FIG. 26, the computing device can acquire a data set (S1038). Further, the computing device can acquire an improved data image based on the acquired data set (S1039). At this time, since the operation S1039 can directly apply the above technical features (FIGS. 16 to 22), a detailed description thereof is omitted. Optionally, the computing device can acquire an improved data set based on the data set (S1040). At this time, since the operation S1040 can directly apply the above technical features (FIGS. 23 to 24), a detailed description thereof is omitted. Further, the computing device can acquire the characteristics of the improved data set based on the data set (S1041). At this time, since the operation S1041 can directly apply the above technical features (FIGS. 10 to 15), it is omitted. Further, the computing device can provide the quality of the data set based on at least one of the data image and the characteristics of the data set (S1042). At this time, for the achievable quality of the data set provided by the computing device, when the data is improved, the method of providing the quality according to the above operation S1037 may be directly applied as it is.
[0265] Computing devices according to various embodiments of the present disclosure can provide a diagnostic report based on various information (e.g., data images, characteristics, improved data images, data set quality, etc.) related to the data set obtained by processing the data set. Specifically, the computing device can provide a comprehensive diagnostic result for the data set via the diagnostic report. In this case, the computing device can output the diagnostic report via an output device (e.g., a display) included in the computing device or an output device of a device communicable with the computing device. For example, when the output device is a display, the computing device can output the diagnostic report to the display screen. Also, for example, when the output device is a VR device, the computing device can output the diagnostic report to a virtual space sent by the VR device.
[0266] FIG. 27 is a diagram showing information included in a diagnostic report provided by a computing device according to various embodiments of the present disclosure.
[0267] FIG. 28 is a diagram showing an example of information for a data image provided by a computing device according to various embodiments of the present disclosure.
[0268] Referring to FIG. 27, the diagnostic report provided by the computing device can include various information about the data set. Specifically, the computing device can provide a diagnostic report including information about the data image, information about the data characteristics, information about the data improvement, and information about the data quality.
[0269] At this time, the information for the data image can include the data image for the data set and the improved data image for the data set. Further, the computing device can provide a diagnostic report further including additional information related to the data image and the improved data image.
[0270] As a specific example, referring to FIG. 28, the diagnostic report provided by the computing device can include information for the data image including the data image IOD or the improved data image MIOD that appears in the imaging space 2800. At this time, the data image IOD can include the point data set 2810 confirmed by finding the manifold where the data set exists. In this case, the computing device can provide the data image IOD by removing the noise of the point data set 2810 according to the description of FIG. 8. Further, the improved data image MIOD can include the improved point data set 2850 obtained by processing the improvement algorithm to obtain the point data set. In this case, the computing device can provide the improved data image MIOD by removing the noise of the improved point data set 2850 according to the description of FIG. 8.
[0271] Also, the diagnostic report provided by the computing device can include additional information related to the data image IOD or the improved data image MIOD. More specifically, the computing device can confirm the point data set by processing the data set to find the manifold where the data set exists, and can obtain various additional information for the data set based on the point data set. Further, the computing device can provide the various additional information obtained as described above together with the data image IOD or the improved data image MIOD.
[0272] The computing device can provide marker information. At this time, the marker information can include markers for areas specified according to preset criteria in the data image IOD or the improved data image MIOD.
[0273] Specifically, the computing device can select a specific area that meets the preset criteria from the data image IOD or the improved data image MIOD, and generate a marker for the area corresponding to the specific area. In this case, the computing device can select the specific area by checking whether the characteristics of the data set meet the preset criteria.
[0274] In one example, the computing device can provide the marker information by generating a marker corresponding to a blank area where there is no data in the data image IOD or the improved data image MIOD. As a specific example, the computing device can provide marker information by generating a marker 2811 for the blank area of the point data set 2810 included in the data image IOD.
[0275] As another example, the computing device can provide the marker information by generating a marker corresponding to a dense area where the data is dense in the data image IOD or the improved data image MIOD.
[0276] As yet another example, the computing device can provide the marker information by generating a marker corresponding to a special area where the distribution of the data is special in the data image IOD or the improved data image MIOD.
[0277] At this time, the computing device can determine an area for generating markers in the point data set 2810 or the improved point data set 2850 based on a preset algorithm. For example, the computing device can obtain a feature map for the existence position of point data through a convolution operation based on a pre-stored kernel (see the description for FIG. 15), and can determine the above-mentioned blank area, dense area, or special area based on the feature map.
[0278] In addition, the computing device can provide the marker information by generating at least one marker based on an input received from the outside. Specifically, when the computing device receives a marker generation input for a specific area on the data image IOD or the improved data image MIOD, a marker can be generated in the specific area.
[0279] In addition, when the computing device receives an input for selecting at least one marker from the outside, the computing device can provide enlarged image information that enlarges and shows the distribution of point data in the area corresponding to the at least one marker on the data image IOD or the improved data image MIOD. For example, when the computing device receives an input for selecting the first marker 2813 from the user, the computing device can enlarge the distribution of point data in the area corresponding to the first marker 2813 and provide the first enlarged image 2815, but it is not limited thereto.
[0280] Also, according to an embodiment, when the computing device receives an input for selecting at least one marker generated in the improved image data (MIOD) from the outside, the computing device not only provides enlarged image information showing the distribution of point data in the region corresponding to the at least one marker on the improved data image MIOD, but also provides enlarged image information showing the distribution of point data in the same region as the region corresponding to the at least one marker on the data image IOD. For example, when the computing device receives an input for selecting a second marker 2853 from a user, the computing device can provide both a second enlarged image 2855 showing the distribution of point data in the region corresponding to the second marker 2853 and a first enlarged image 2815 for the same region (e.g., the region where the first marker 2813 is displayed) on the data image IOD, but is not limited thereto.
[0281] In addition, the computing device can provide manifold boundary information 2817 by displaying the manifold boundary of the data image IOD or the improved data image MIOD. Specifically, the computing device can provide the manifold boundary information 2817 by displaying the boundary region of the manifold formed by the point data set 2810 confirmed based on the data set.
[0282] In addition, the computing device can provide grouping information (not shown) for the data image IOD or the improved data image MIOD. Specifically, when the point data included in the point data set 2810 or the improved point data set 2850 is clustered into at least one or more groups, the computing device can provide the grouping information by adding a display showing the clustered point data.
[0283] In addition, the computing device can add a visual effect to the data image IOD or the improved data image MIOD. Specifically, the computing device can represent the point data included in the point data set 2810 or the improved point data set 2850 using a preset color or shape in order to enhance the visual effect of the data image IOD or the improved data image MIOD. For example, the computing device can represent the color of the point data included in the dense region of the data to be different from the color of other point data in order to indicate the density of the data set, but is not limited thereto. Also, for example, the computing device can represent the point data clustered into different groups using different shapes, but is not limited thereto.
[0284] In addition, the computing device can provide comparison information (not shown) indicating the difference between the data image IOD and the improved data image MIOD. Specifically, the computing device can display the changed part in the improved data image MIOD compared with the existing data image IOD by improving the data set. For example, since the computing device generates the improved point data set 2850 based on the point data set 2810, it can display the region where the distribution of the point data has changed on the improved point data set 2850 based on the point data set 2810, but is not limited thereto.
[0285] Referring back to FIG. 27, the computing device can provide a diagnostic report including information on data characteristics. At this time, the information on the data characteristics can include the characteristics of the acquired data set and the characteristics of the improved data set, but is not limited thereto, and can further include additional information that can be obtained based on the characteristics of the data set and the characteristics of the improved data set.
[0286] In addition, the computing device can provide a diagnostic report including information on data improvement. At this time, the information on data improvement can include, but is not limited to, an improved data set, and can further include additional information obtainable based on the improved data set. For example, the information on data improvement can include virtual data generated based on improved point data included in the improved data set. Also, for example, the information on data improvement can include sample information obtained by extracting a part of the virtual data.
[0287] In addition, the computing device can provide a diagnostic report including information on data quality. At this time, the information on data quality can include, but is not limited to, the quality of the acquired data set and the achievable quality of the data set, and can further include additional information obtainable based on the quality of the data set and the achievable quality of the data set.
[0288] FIG. 29 is a diagram for explaining the operation of a computing device according to various embodiments of the present disclosure to provide a data image and an improved data image of a data set.
[0289] Referring to FIG. 29, the computing device can acquire a data set (S1043). Also, the computing device can confirm a first point data set by mapping the acquired data set to a first embedding space (S1044). At this time, since the operations of S1043 and S1044 can directly apply the above technical features (FIGS. 4 to 9), detailed description thereof is omitted.
[0290] Further, the computing device can identify a second point data set by mapping the identified first point data set to a second embedding space (S1045). Further, the computing device can identify an improved first point data set by restoring the identified second point data set to the first embedding space (S1046). At this time, since the operations of S1045 and S1046 can directly apply the above technical features (FIGS. 16 to 22), detailed descriptions thereof are omitted.
[0291] Further, the computing device can provide a data image based on the first point data set and provide an improved data image based on the improved first point data set (S1047). At this time, since the specific method by which the computing device provides a data image based on the first point data set can directly apply the above technical features (FIGS. 4 to 9), it is omitted. Also, in the operation of S1047, the computing device can represent the improved data image in the same imaging space as the data image. Also, without being limited thereto, the computing device can represent the improved data image in an imaging space different from the data image.
[0292] Alternatively or additionally, the computing device can obtain the characteristics of the data set based on the first point data set, and can obtain the improved characteristics of the data set (modified property) based on the improved first point data set. At this time, since the specific method of obtaining the characteristics of the data set can directly apply the above technical features (FIGS. 10 to 15), it is omitted.
[0293] Alternatively or additionally, the computing device can provide an improved data set including virtual data by restoring the improved point data set to an output domain. At this time, the specific method of providing the improved data set can directly apply the above technical features (Figs. 23 to 24), so it is omitted here.
[0294] FIG. 30 is a diagram showing an algorithm execution model constituting a computing device according to various embodiments of the present disclosure.
[0295] Referring to FIG. 30, the computing device 3000 can include a plurality of algorithm execution models with different purposes. Specifically, the computing device 3000 can include a plurality of algorithm execution models designed to output a specific output. For example, the computing device can include, but is not limited to, an imaging model 3100 designed to provide a data image, an improvement model 3200 designed to provide improved point data, a generation model 3300 designed to generate an improved data set including virtual data, a feature extraction model 3400 designed to calculate the characteristics of data, and a diagnosis model 3500 designed to provide a diagnosis report. Of course, the plurality of algorithm execution models may be configured by one integrated model.
[0296] In addition, the computing device can selectively output output data by selectively inputting input data to at least a part of the plurality of algorithm execution models. At this time, the computing device can determine which model to process the data set based on the user input input together with the data set. For example, when the computing device obtains a data set together with user input so as to output a data image, the computing device can output a data image by inputting the data set to the imaging model 3100.
[0297] Also, among the plurality of algorithm models, the output data of a specific model can be used as the input data of other models. For example, when a computing device acquires a data set together with user input so as to generate virtual data, the computing device can acquire improved point data obtained by inputting the data set into the improvement model 3200, and the improved point data set can be input into the generation model 3300 to provide an improved point data set including virtual data.
[0298] Also, for example, when a computing device acquires a data set together with user input so as to generate an improved data image, the computing device can acquire improved point data obtained by inputting the data set into the improvement model 3200, and the improved point data set can be input into the imaging model 3100 to provide an improved data image.
[0299] Also, for example, when a computing device acquires a data set together with user input so as to generate a diagnostic report, the computing device can provide a diagnostic report by inputting the data image, the improved data image, the improved point data set, the characteristics of the data set, and the improved characteristics of the data set acquired based on the data set into the diagnostic model 3500.
[0300] FIG. 31 is a diagram showing a method in which at least one processor included in a computing device according to various embodiments of the present disclosure selectively executes operations based on a data set.
[0301] Referring to FIG. 31, the at least one processor can obtain a data set (S1048). Further, the at least one processor can determine the data set in a preset manner (S1049). For example, the at least one processor can determine the capacity, application domain, modality, type, or number of modalities of the data set.
[0302] In addition, the at least one processor can determine the data set based on a pre-stored algorithm. Also, the at least one processor can determine the data set by searching for data similar to the obtained data set in a pre-stored database.
[0303] Further, the at least one processor can execute an operation based on at least one of a plurality of instructions stored in the memory of the computing device according to the determination result (S1050).
[0304] Specifically, the at least one processor can execute a process instructed by at least one instruction determined based on a confirmed trigger as a result of determining the data set. At this time, the trigger can be an event that triggers the operation of the at least one processor, and the process executed by the at least one processor can be determined according to the type of the trigger. More specifically, the trigger can be, but is not limited to, an event that instructs to provide specific output data.
[0305] A specific example will be described with reference to FIG. 32.
[0306] FIG. 32 is a diagram showing various processes executed by at least one processor according to instructions stored in the memory of a computing device according to various embodiments of the present disclosure.
[0307] Referring to FIG. 32, at least one processor of the computing device can operate based on one of a plurality of processes (data processing pipelines) according to the trigger when the trigger is confirmed.
[0308] Specifically, when a first trigger occurs, at least one processor can operate based on a first process 3210. At this time, the at least one processor can operate based on at least a part of a plurality of instructions included in the first process 3210.
[0309] For example, when the first trigger instructs to provide a data image, the at least one processor is instructed by an instruction 3211 to perform an operation of checking a point data set based on a data set acquired by the at least one processor, and an instruction 3213 to perform an operation of providing a data image based on the data set, and can operate based thereon. Of course, the at least one processor may further perform an operation based on an instruction 3213 to perform an operation of acquiring characteristics of the data set based on the point data set.
[0310] Also, for example, when the first trigger instructs to provide characteristics of a data set, the at least one processor can operate based on an instruction 3211 to perform an operation of checking a point data set based on a data set acquired by the at least one processor, and an instruction 3213 to perform an operation of acquiring characteristics of the data set based on the point data set.
[0311] In addition, the computing device can pre-store information about the first trigger connected to the first process 3210. Specifically, the first trigger can include receiving user input instructing to provide a data image and a judgment result for the data set. Also, the first trigger may occur immediately when the data set is input. In other words, the first trigger instructing to provide a data image can be a basic trigger that occurs simultaneously with the acquisition of the data set, but is not limited thereto.
[0312] Also, when a second trigger occurs, at least one processor can operate based on the second process 3220. At this time, the at least one processor can operate based on at least a part of a plurality of instructions included in the second process 3220.
[0313] For example, when the second trigger instructs to provide an improved point data set, the at least one processor executes an instruction 3221 to instruct to perform an operation of checking the point data set based on the data set acquired by the at least one processor and an instruction 3222 to instruct to perform an operation of checking the improved point data set based on the point data set by the at least one processor. can operate based on
[0314] Also, for example, when the second trigger is instructed to provide improved characteristics of a dataset, the at least one processor can operate based on instruction 3221 that instructs the at least one processor to perform an operation of verifying a point dataset based on the dataset acquired by the at least one processor, instruction 3222 that instructs the at least one processor to perform an operation of verifying an improved point dataset based on the point dataset, and instruction 3223 that instructs the at least one processor to perform an operation of acquiring improved characteristics of the dataset based on the improved point dataset.
[0315] Also, for example, when the second trigger is instructed to provide an improved data image of a dataset, the at least one processor can operate based on instruction 3221 that instructs the at least one processor to perform an operation of verifying a point dataset based on the dataset acquired by the at least one processor, instruction 3222 that instructs the at least one processor to perform an operation of verifying an improved point dataset based on the point dataset, and instruction 3224 that instructs the at least one processor to perform an operation of providing an improved data image based on the improved point dataset.
[0316] Also, the computing device can pre-store information about the second trigger connected to the second process 3220. Specifically, the second trigger can include reception of user input instructing to provide an improved data image and a judgment result for the dataset.
[0317] Also, when a third trigger occurs, the at least one processor can operate based on a third process 3230. At this time, the at least one processor can operate based on at least a part of a plurality of instructions included in the third process 3230.
[0318] For example, when the third trigger instructs to provide the quality of a data set, the at least one processor can operate based on an instruction 3231 that instructs the at least one processor to perform an operation of verifying a point data set based on the data set acquired by the at least one processor, and an instruction 3233 that instructs the at least one processor to perform an operation of acquiring the quality of the data set based on the point data set.
[0319] Also, for example, when the third trigger instructs to provide the achievable quality of a data set, the at least one processor can operate based on an instruction 3231 that instructs the at least one processor to perform an operation of verifying a point data set based on the data set acquired by the at least one processor, an instruction 3232 that instructs the at least one processor to perform an operation of verifying an improved point data set based on the point data set, and an instruction 3234 that instructs the at least one processor to perform an operation of acquiring the achievable quality of the data set based on the improved point data set.
[0320] Also, the computing device can pre-store information for the third trigger connected to the third process 3230. Specifically, the third trigger can include receiving a user input instructing to provide the achievable quality of a data set and a judgment result for the data set.
[0321] The selective operation of the at least one processor is not limited to the process shown in FIG. 32, and can selectively execute the operation of the processor according to a trigger generated based on the output that can be output by a computing device according to various embodiments of the present disclosure. For example, when a fourth trigger (not shown) instructs to provide a diagnostic report, the at least one processor can operate based on at least one instruction that instructs to obtain information necessary to generate the diagnostic report.
[0322] In addition, a computing device according to various embodiments can configure a preset database by databaseizing a plurality of processes consisting of a plurality of instructions as described above. Specifically, the computing device can store all of the above methods (for example, data imaging, feature extraction, improvement, evaluation, etc.), input data and output data associated with the method, and further, a method for generating a manifold associated with the method (for example, a dimension determination method, an optimized form determination method, etc.) to configure a preset database.
[0323] In addition, when a data set is input, the computing device can select at least one of the plurality of processes stored in the preset database, and can process the data set based on the selected process.
[0324] In addition, the computing device can reconfigure the preset database. More specifically, the computing device can repeatedly execute an optimization process so as to generate a more optimized output, rather than processing the input data set according to the initially determined process to generate a final output, whereby the preset database can be reconfigured based on the optimized process. For example, the computing device can reconfigure the preset database based on a machine learning method, but is not limited thereto.
[0325] FIG. 33 is a diagram showing a configuration example of a computing device according to various embodiments of the present disclosure.
[0326] Referring to FIG. 33, the computing device can include various configurations for outputting various output data based on a data set defined on an input domain.
[0327] Specifically, the computing device can include a first converter 3310 designed to generate a first manifold based on the acquired data set. At this time, the first manifold can be defined on a first embedding space. Also, the first converter 3310 can convert the data set into the first manifold based on a first preset function. Further, the computing device can include a second converter 3330 designed to generate a second manifold based on the first manifold. At this time, the second manifold can be defined on a second embedding space having a different dimension from the first embedding space. Also, the second converter 3330 can convert the first manifold into the second manifold based on a second preset function. For example, the first converter 3310 and the second converter 3330 can include, but are not limited to, an encoder.
[0328] Further, the computing device may include a first restorer 3320 designed to generate first restored data based on the first manifold. At this time, the first restored data can be defined on an output domain having the same dimension as the input domain. Also, the first restorer 3320 can restore the first manifold with the first restored data based on the inverse function of the first preset function. Further, the computing device may include a second restorer 3340 designed to generate an improved first manifold based on the second manifold. At this time, the improved first manifold can be defined on a third embedding space having the same dimension as the first embedding space. Also, the second restorer 3340 can restore the second manifold with the improved first manifold based on the inverse function of the second preset function.
[0329] Further, the computing device may include a feature extractor 3350 designed to generate characteristics of a data set based on the first manifold and generate improved characteristics of the data set based on the second manifold or the improved first manifold. At this time, the characteristics of the data set or the improved characteristics of the data set can be provided in the form of a feature map. Also, the feature extractor 3350 may be provided in the form of a feed-forward neural network.
[0330] In addition, the computing device can include an imaging device 3360 designed to generate a data image based on the first manifold and generate an improved data image based on the improved first manifold. At this time, the data image and the improved data image can appear on a predefined imaging space. Further, the imaging device 3360 can represent the first manifold and the improved first manifold as the data image and the improved data image respectively based on a preset data visualization algorithm.
[0331] The method according to the embodiment can be realized in a program instruction form executable through various computer means and recorded on a computer-readable medium. The computer-readable medium can include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded on the medium may be specially designed and configured for the embodiment, or may be known and usable by those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions such as ROM, RAM, and flash memory. Examples of program instructions include not only machine language code created by a compiler but also high-level language code that can be executed by a computer using an interpreter or the like. The above hardware device can be configured to operate as one or more software modules for executing the operations of the embodiment, and vice versa.
[0332] Although the embodiments have been described by way of limited embodiments and drawings as above, various modifications and variations are possible for those having ordinary knowledge in the relevant technical field from the above description. For example, the described technology may be executed in an order different from the described method, and / or the components such as the described system, structure, device, circuit, etc. may be combined or combined in a form different from the described method, or replaced or substituted by other components or equivalents, and appropriate results can be achieved.
[0333] Therefore, other configurations, other embodiments, and equivalents to the claims, etc. also fall within the scope of the claims described below.
Claims
1. In a method of operating at least one processor included in a computing device, the step of obtaining a data set; the step of identifying a first point data set including respective point data corresponding to respective data included in the data set by mapping the data set into a first embedding space; the step of providing a data image (Image of Data, IOD) obtained by displaying the first point data set in an imaging space; the step of obtaining an improved first point data set including at least one improved point data not included in the first point data set based on the first point data set; the step of providing an improved data image (Modified Image of Data, MIOD) obtained by displaying the improved first point data set in the imaging space, the method comprising these steps.
2. The step of identifying the first point data set includes the step of identifying a first manifold generated by mapping the data set onto the first embedding space according to a first preset condition, and the step of identifying the first point data set included in the first manifold, wherein the first manifold includes the shape formed by the first point data set on the first embedding space, The method according to Claim 1.
3. The step of identifying the first point data set further includes the step of providing a restored data set by restoring the first point data set, wherein the restored data set has a modality corresponding to the data set, and the first preset condition is set based on the similarity between the data set and the restored data set. The method according to Claim 2.
4. The method further includes the step of identifying a second point data set by mapping the first point data set into a second embedding space, wherein the improved first point data set is obtained by restoring the second point data set into the first embedding space. The method according to Claim 1.
5. The step of identifying the second point data set includes Confirming a second manifold generated by mapping the first point data set onto a second embedding space according to a second preset condition; confirming the second point data set included in the second manifold, including: The second preset condition is set based on the similarity between points included in the first point data set. The method according to claim 4.
6. The at least one improved point data is generated by restoring at least one point data included in the second point data set to the first embedding space. The method according to claim 5.
7. The step of obtaining the improved first point data set includes: clustering the first point data set into at least one group; adjusting the distance on the first embedding space between the first point data included in the first group and the second point data included in a second group different from the first group. The method according to claim 1.
8. The distance between the first point data and the second point data is farther than the distance between the first point data and third point data, wherein the third point data is included in the first group. The method according to claim 7.
9. The step of providing the data image includes: confirming, on the first embedding space, a manifold region formed by the first point data set; obtaining the data image by representing the first point data set in the imaging space such that at least one point data located outside the manifold region is deleted. The method according to claim 1.
10. The imaging space is a space in which the data image and the improved data image are sent by at least one output device connected to the computing device. The method according to claim 1.
11. A first transformer designed to generate a first manifold based on a dataset defined on an input domain, where the first manifold is defined on a first embedding space and includes a first point dataset corresponding to the dataset. A first restorer designed to generate an improved first manifold based on the first manifold, where the improved first manifold includes at least one improved point data not included in the first point dataset. An imaging unit designed to provide a data image by representing the first manifold in an imaging space and provide an improved data image by representing the improved first manifold in the imaging space. A computing device. **Claim 12** Further comprising a second transformer designed to generate a second manifold based on the first manifold. At this time, the second manifold is defined on a second embedding space having a different dimension from the first embedding space and includes a second point dataset corresponding to the first point dataset. The first restorer is designed to generate the improved first manifold by restoring the second manifold to the first embedding space. The computing device according to claim 11. **Claim 13** Further comprising a second restorer designed to generate first restored data based on the first manifold. At this time, the first restored data is defined on an output domain having the same dimension as the input domain. The computing device according to claim 11. **Claim 14** The second restorer is designed to generate second restored data based on the improved first manifold. The second restored data includes virtual data generated based on at least one improved point included in the improved first manifold. The computing device according to claim 13. **Claim 15** Further comprising a feature extractor designed to extract characteristics of a dataset based on the first manifold and extract improved characteristics of the dataset based on the improved first manifold. The computing device according to claim 11.
Citation Information
Patent Citations
Methods and systems for reducing dimensionality in a reduction and prediction framework
US20210241113A1
Data extension device, learning device, data extension method, and recording medium
WO2022009254A1