Data preprocessing device
The data preprocessing device automates data preprocessing using machine learning models to enhance data quality and speed up AI development, addressing the inefficiencies of manual methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- D4ALL CO LTD
- Filing Date
- 2025-11-11
- Publication Date
- 2026-05-07
AI Technical Summary
Existing data preprocessing methods in AI development are slow and labor-intensive, hindering the efficiency of AI and machine learning processes.
A data preprocessing device that utilizes two machine learning models to automate data preprocessing by generating models for context identification and processing identification, followed by format conversion, enhancing data quality and enabling faster AI development.
The device accelerates AI development and reduces labor costs by automating data preprocessing, improving data quality and model performance through processes like imputing missing values, removing inappropriate data, and normalizing data.
Smart Images

Figure 0007854759000001_ABST
Abstract
Description
Technical Field
[0001] It relates to the technology of data science or AI (Artificial Intelligence) development.
Background Art
[0002] In recent years, the information processing ability of computers has been significantly improved. Thanks to this, the technologies of data science and AI that analyze and process a large amount of information and produce valuable information have attracted attention.
[0003] Under such a background, technical proposals in the fields of data science and AI have been actively made. For example, in Patent Document 1, an information processing system that relaxes the expertise in data science required of users when converting unstructured data into structured data has been proposed.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Here, preprocessing refers to the process of converting data into a form that can be utilized in data science or AI development, and is an important step for enabling AI (Artificial Intelligence) and machine learning models to learn data more accurately.
[0006] Therefore, an object of the present invention is to provide a data preprocessing device that accelerates the speed of AI development and achieves labor saving by automating data preprocessing in AI development.
Means for Solving the Problems
[0007] One form of the data preprocessing device disclosed includes: a first model generation means that trains a training dataset consisting of a combination of management information and contextual information relating to the management information, and generates a first machine learning model that outputs contextual information relating to the management information when the management information is input; a second model generation means that trains a training dataset consisting of a combination of the management information, contextual information relating to the management information, and the content of preprocessing to be performed on the management information, and generates a second machine learning model that outputs the content of preprocessing to be performed on the management information when the management information and the contextual information relating to the management information are input; and the first machine learning model and the second machine learning model The system comprises: a model storage means for storing parameters that define the operation of each; a context identification means for inputting one of the management information to the first machine learning model and identifying the output of the first machine learning model as one of the context information relating to the one of the management information; a processing identification means for inputting the one of the management information and the one of the context information to the second machine learning model and identifying the output of the second machine learning model as the content of the preprocessing to be performed on the one of the management information; and an information format conversion means for converting the one of the management information according to the content of the preprocessing to be performed on the identified one of the management information, wherein the preprocessing is a process that converts the target data into data in a format that can be used for data science or AI (Artificial Intelligence) development. [Effects of the Invention]
[0008] The data preprocessing device being disclosed will accelerate the speed of AI development and reduce labor costs by automating data preprocessing in AI development. [Brief explanation of the drawing]
[0009] [Figure 1] This figure illustrates the overview of the data preprocessing device according to this embodiment. [Figure 2] This is a functional block diagram of the data preprocessor according to this embodiment. [Figure 3]This figure shows an example of the hardware configuration of the data preprocessing device according to this embodiment. [Figure 4] This flowchart shows an example of the processing flow by the data preprocessing device according to this embodiment. [Modes for carrying out the invention]
[0010] The embodiments for carrying out the present invention will be described with reference to the drawings. (Operating principle of the data preprocessor according to this embodiment)
[0011] The operating principle of the data preprocessing device (hereinafter simply referred to as "this device") 100 according to this embodiment will be explained using Figures 1 and 2. Figure 1 is a diagram showing the connection relationship between this device 100 and other devices, and Figure 2 is a functional block diagram of this device 100.
[0012] As shown in Figure 1, the device 100 is connected to one or more external devices 260 via a communication network 270. The communication network 270 may be a wired communication network or a wireless communication network. The external devices 260 may be, for example, a POS (Point of Sales) system.
[0013] As shown in Figure 2, the device 100 includes a pre-processing information storage means 110, a post-processing information storage means 120, a model storage means 130, a first model generation means 140, a second model generation means 150, a context identification means 160, a processing identification means 170, and an information format conversion means 180.
[0014] The pre-processing information storage means 110 stores the management information 210 before processing by the information format conversion means 180. The management information 210 may be, for example, customer purchase information, skin information, health management information, etc., but it may be any other type of information. The management information 210 is provided, for example, from a POS system 260.
[0015] The post-processing information storage means 120 stores the management information 215 after processing by the information format conversion means 180. The management information 215 is data in a format that can be used for data science or AI (Artificial Intelligence) development. The management information 215 is, for example, qualitative data, quantitative data, or annotation data, where qualitative data is text data and quantitative data is data in the form of discrete or continuous values.
[0016] Furthermore, the management information 215 is data in an annotated format in which the management information 210 is vectorized, aggregated, written, or annotated, and aggregation means converting it into descriptive statistics or converting it into a grand total.
[0017] The model storage means 130 stores parameters that define the operation of the trained machine learning models 240 and 250, which have been trained by the first model generation means 140 and the second model generation means 150, which will be described later.
[0018] The first model generation means 140 trains the machine learning model 240 with a training dataset consisting of combinations of management information 210 and context information 220 related to the management information 210. The training dataset consisting of combinations of management information 210 and context information 220 related to the management information 210 is prepared in advance. By doing so, the first model generation means 140 generates a first machine learning model 240 that, upon input of management information 210, outputs context information 220 related to the management information 210. The learning algorithm used by the first model generation means 140 is not particularly limited.
[0019] Here, the context information 220 is background information that determines what meaning the target data has. The context information 220 includes, for example, time information (time stamps (creation date, update date), temporal position (past / present / future prediction), temporal urgency (immediacy, real-time), expiration date, shelf life, seasonality, periodicity... etc.), location / space information (geographical location (country, region, city), physical / virtual space, geopolitical risk area, regulatory jurisdiction area, time zone... etc.), subject information (attributes of the information sender (organization, individual, position), target user group (expert / general, age group), stakeholder relationship, authority level, security clearance, reliability / credibility score... etc.), purpose / intention information (business purpose (profit / non-profit), usage purpose (analysis / reporting / decision-making), strategic importance, risk management purpose, compliance requirements... etc.), content information (type of information (numerical / text / image), sensitivity of data (confidentiality level), topic / domain, data quality / completeness, version information... etc.), method information (data acquisition method (automatic / manual), processing history, conversion history, usage technology / protocol, data format, access method... etc.), other external factor information (laws and regulations / compliance (GDPR, personal information protection law, industry regulations (finance, medical, etc.), export control regulations, copyright / intellectual property rights... etc.), geopolitical factors (international relations / diplomatic situation, countries subject to economic sanctions, data localization requirements, security considerations... etc.), market / business environment (competitive situation, market trends, economic indicators, industry standards... etc.), technical constraints (system environment, network status, device type, browser / OS version... etc.), other internal factor information (organizational factors (internal company policies, governance system, organizational hierarchy, approval process... etc.), data quality factors (data freshness, accuracy / correctness, completeness, consistency... etc.), security factors (access rights, encryption requirements, audit log requirements, incident history... etc.), user behavior factors (past operation history, reference pattern, error occurrence tendency, learning progress... etc.), results of information processing performed on the management information of the target, metadata (source / reference)).
[0020] The second model generation means 150 causes the machine learning model 240 to learn from a learning data set composed of a combination of the management information 210, the context information 220 regarding the management information 210, and the content 230 of the preprocessing to be performed on the management information 210. The learning data set composed of the combination of the management information 210, the context information 220 regarding the management information 210, and the content 230 of the preprocessing to be performed on the management information 210 is prepared in advance. By doing so, the second model generation means 150 generates a second machine learning model 250 that outputs the content 230 of the preprocessing to be performed on the management information 210 when the management information 210 and the context information 220 regarding the management information 210 are input. Note that the learning algorithm by the second model generation means 150 is not particularly limited.
[0021] The context identification means 160 inputs a piece of management information 210 to the first machine learning model 240 and causes the first machine learning model 240 to output a piece of context information 220 regarding the piece of management information 210. The context identification means 160 recognizes (identifies) the output of the first machine learning model 240 as a piece of context information 220 regarding a piece of management information 210.
[0022] The process identification means 170 inputs a piece of management information 210 and a piece of context information 220 regarding the piece of management information 210 to the second machine learning model 250 and causes the second machine learning model 250 to output the content 230 of the preprocessing to be performed on the piece of management information 210. The process identification means 150 recognizes (identifies) the output of the second machine learning model 250 as the content 230 of the preprocessing to be performed on a piece of management information 210.
[0023] The information format conversion means 180 converts a piece of management information 210 identified by the processing identification means 170 according to the content of the preprocessing 230 to be applied to that piece of management information 210. The preprocessing 230 is a process to convert the target data 210 into data in a format that can be used for data science or AI development. Preprocessing 230 is an important step to enable AI (artificial intelligence) and machine learning models to learn the data more accurately, and specifically includes imputing missing values in the data, removing inappropriate data, normalizing the data, and selecting features. Through these operations, the quality of the data can be improved and the performance of the model can be maximized.
[0024] Here, the preprocessing content 230 is a process of converting management information 210 into qualitative data, quantitative data, or annotation data, where qualitative data is text data and quantitative data is discrete or continuous data.
[0025] Furthermore, the preprocessing content 230 refers to the process of vectorizing, aggregating, writing text, or annotating the management information 210, and aggregation means converting to descriptive statistics or converting to a grand total.
[0026] Based on the operating principle described above, this device 100 automates data preprocessing in AI development, thereby accelerating the speed of AI development and saving labor. (Hardware configuration of the data preprocessing device according to this embodiment)
[0027] An example of the hardware configuration of the device 100 will be explained using Figure 3. Figure 3 is a diagram showing an example of the hardware configuration of the device 100. As shown in Figure 3, the device 100 has a CPU (Central Processing Unit) 510, ROM (Read-Only Memory) 520, RAM (Random Access Memory) 530, auxiliary storage device 540, communication I / F 550, input device 560, display device 570, and storage medium I / F 580.
[0028] The CPU 510 is a device that executes programs stored in the ROM 520. It processes data loaded into the RAM 530 according to program instructions and controls the entire device 100. The ROM 520 stores the programs and data that the CPU 510 will execute. When the CPU 510 executes a program stored in the ROM 520, the RAM 530 loads the programs and data to be executed and temporarily holds the calculation data during the calculation.
[0029] The auxiliary storage device 540 is a device that stores the operating system (OS), which is the basic software, and the application programs according to this embodiment, along with related data. For example, it may include a pre-processing information storage means 110, a post-processing information storage means 120, and a model storage means 130. The auxiliary storage device 540 is, for example, an HDD (Hard Disk Drive) or flash memory.
[0030] The communication interface 550 is an interface for exchanging data with other devices (such as POS systems) 260 that provide communication functions, by connecting to a communication network 270 such as a wired or wireless LAN (Local Area Network) or the Internet.
[0031] The input device 560 is a device for inputting data into the main device 100, such as a keyboard. The display device (output device) 570 is a device consisting of an LCD (Liquid Crystal Display) or the like, and functions as a user interface for the user to use the functions of the main device 100 and to make various settings. The storage medium I / F 580 is an interface for sending and receiving data with storage media 590 such as CD-ROMs, DVD-ROMs, and USB memory.
[0032] Each of the means of this device 100 may be realized by the CPU 510 executing a program corresponding to each means stored in the ROM 520 or auxiliary storage device 540. Alternatively, each of the means of this device 100 may be realized by the processing related to each means being implemented as hardware. Furthermore, the program according to the present invention may be read from an external server device via a communication I / F 550, or read from a storage medium 590 via a storage medium I / F 580, and the device 100 may execute the program. (Example of processing by the data preprocessing device according to this embodiment) The flow of a processing example using this device 100 will be explained using Figure 4. Figure 4 is a flowchart showing the flow of a processing example using this device 100.
[0033] In S10, the first model generation means 140 trains the machine learning model 240 with a training dataset consisting of a combination of management information 210 and context information 220 related to the management information 210. By doing so, the first model generation means 140 generates a first machine learning model 240 that, when it receives management information 210 as input, outputs context information 220 related to the management information 210. Here, the context information 220 is background information that determines what meaning the target data has.
[0034] Furthermore, in S10, the second model generation means 150 trains the machine learning model 240 with a training dataset consisting of a combination of management information 210, context information 220 related to management information 210, and the content of preprocessing to be applied to management information 210 230. By doing so, the second model generation means 150 generates a second machine learning model 250 that, upon input of management information 210 and context information 220 related to management information 210, outputs the content of preprocessing to be applied to management information 210 230.
[0035] Here, the preprocessing content 230 is a process of converting management information 210 into qualitative data, quantitative data, or annotation data, where qualitative data is text data and quantitative data is discrete or continuous data.
[0036] Furthermore, the preprocessing content 230 refers to the process of vectorizing, aggregating, writing text, or annotating the management information 210, and aggregation means converting to descriptive statistics or converting to a grand total.
[0037] In S20, the context identification means 160 inputs a piece of management information 210 to the first machine learning model 240 and causes the first machine learning model 240 to output a piece of context information 220 relating to the piece of management information 210. The context identification means 160 recognizes (identifies) the output of the first machine learning model 240 as a piece of context information 220 relating to the piece of management information 210.
[0038] In S30, the processing identification means 170 inputs one piece of management information 210 and one piece of context information 220 identified in S20 to the second machine learning model 250, causing the second machine learning model 250 to output the content 230 of the preprocessing to be applied to the one piece of management information 210. The processing identification means 150 recognizes (identifies) the output of the second machine learning model 250 as the content 230 of the preprocessing to be applied to the one piece of management information 210.
[0039] In S40, the information format conversion means 180 converts a management information 210 according to the content of the preprocessing 230 to be performed on the management information 210 identified in S30. The preprocessing 230 is a process to convert the target data 210 into data in a format that can be used for data science or AI development. Preprocessing 230 is an important step to enable AI (artificial intelligence) and machine learning models to learn the data more accurately, and specifically includes imputing missing values in the data, removing inappropriate data, normalizing the data, and selecting features. Through these operations, the quality of the data can be improved and the performance of the model can be maximized.
[0040] By performing the above-described processing, the device 100 automates data preprocessing in AI development, thereby accelerating the speed of AI development and reducing labor costs.
[0041] Although embodiments of the present invention have been described in detail above, the present invention is not limited to these specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention as described in the claims. [Explanation of Symbols]
[0042] 100 Data Preprocessing Devices 110 Pre-processing information storage means 120 Post-processing information storage means 130 Model Storage Methods 140 First Model Generation Means 150 Second Model Generation Means 160 Context Identification Means 170 Processing Identification Means 180 Information format conversion means 210 Management information 215 Management information after processing by information format conversion means 220 Contextual Information 230 Pre-processing 240 First Machine Learning Model 250 Second Machine Learning Model External devices including 260 POS systems 270 Communication Networks 510 CPU 520 ROM 530 RAM 540 Auxiliary storage 550 Communication Interfaces 560 Input Device 570 Output device 580 Storage Media Interface 590 Storage medium
Claims
1. A first model generation means generates a first machine learning model that trains on a training dataset consisting of a combination of management information and contextual information related to said management information, and outputs contextual information related to said management information when said management information is input. A second model generation means generates a second machine learning model that, when the management information and the context information related to the management information are input, outputs the content of the preprocessing to be applied to the management information, by training on a training dataset consisting of a combination of the management information, context information related to the management information, and the content of the preprocessing to be applied to the management information. A model storage means for storing parameters that define the operation of the first machine learning model and the second machine learning model, A context identification means that inputs one piece of the management information to the first machine learning model and identifies the output of the first machine learning model as one piece of the context information relating to the one piece of management information, A processing identification means that inputs the first management information and the first context information to the second machine learning model and identifies the output of the second machine learning model as the content of the preprocessing to be performed on the first management information, The system includes an information format conversion means for converting the specified management information according to the content of the preprocessing to be performed on the specified management information, A data preprocessing device characterized in that the preprocessing is a process of converting target data into data in a format that can be used for data science or AI (Artificial Intelligence) development.
2. The data preprocessing apparatus according to claim 1, characterized in that the preprocessing is a process of converting the management information into annotation data, which is qualitative data, quantitative data, or data to which annotations have been added.
3. The data preprocessing apparatus according to claim 2, characterized in that the qualitative data is text data and the quantitative data is discrete or continuous data.
4. The data preprocessing device according to claim 1, characterized in that the preprocessing is a process of vectorizing, aggregating, textualizing, or annotating the management information.
5. The data preprocessing device according to claim 4, characterized in that the aggregation is a process of converting to a descriptive statistic or a process of converting to a grand total.
6. A data preprocessing method performed by a computer, The first model generation means trains a training dataset consisting of a combination of management information and contextual information related to the management information, and generates a first machine learning model that outputs contextual information related to the management information when the management information is input. The second model generation means trains a training dataset consisting of a combination of the management information, contextual information relating to the management information, and the content of the preprocessing to be applied to the management information, and generates a second machine learning model that outputs the content of the preprocessing to be applied to the management information when the management information and the contextual information relating to the management information are input. The context identification means inputs one piece of the management information to the first machine learning model and identifies the output of the first machine learning model as one piece of the context information relating to the one piece of management information, The processing identification means inputs the first management information and the first context information to the second machine learning model, and identifies the output of the second machine learning model as the content of the preprocessing to be performed on the first management information. The information format conversion means includes the step of converting the specified management information in accordance with the content of the preprocessing to be performed on the specified management information, A data preprocessing method characterized in that the preprocessing is a process of converting target data into data in a format that can be used for data science or AI (Artificial Intelligence) development.
7. The data preprocessing method according to claim 6, characterized in that the preprocessing is a process of converting the management information into annotation data, which is qualitative data, quantitative data, or data to which annotations have been added.
8. The data preprocessing method according to claim 7, characterized in that the qualitative data is text data and the quantitative data is discrete or continuous data.
9. The data preprocessing method according to claim 6, characterized in that the preprocessing is a process of vectorizing, aggregating, textualizing, or annotating the management information.
10. The data preprocessing method according to claim 9, characterized in that the aggregation is a process of converting to a descriptive statistic or a process of converting to a grand total.
11. A data preprocessing program for causing a computer to perform the method according to any one of claims 6 to 10.
Citation Information
Patent Citations
Method for vectorizing medical data for machine learning, data conversion device and data conversion program implementing the method
JP2024522648A
Information processing system, information processing method, and information processing program
JP2024039064A